FAX uses explicit claim verification against faithful tools to raise simulation faithfulness from 0.20 to 0.46 on the new CRAFTER-XAI-Bench open-world RL benchmark.
Towards faithfully interpretable nlp systems: How should we define and evaluate faithfulness?
11 Pith papers cite this work, alongside 3 external citations. Polarity classification is still indexing.
representative citing papers
GAE reduces the faithfulness gap in dictionary-based explainers under distribution shift by geometrically realigning the ID dictionary to the OOD-active subspace, with a quadratic excess-loss bound.
EEG foundation models show no single winner across failure modes, attend to correct brain regions but decode corrupted signals, and retain task information in early layers while late layers adapt during fine-tuning.
Chain-of-thought monitoring detects reward hacking in frontier reasoning models, but strong optimization against the monitor produces obfuscated misbehavior that remains hard to detect.
Introduces the TCR framework to evaluate educational LLM assistants on transparency, consistency, and refinement in multi-turn interactions, complementing aggregate metrics.
A latent mediation framework with sparse autoencoders enables non-additive token-level influence attribution in LLMs by learning orthogonal features and back-propagating attributions.
AtManRL learns an additive attention mask on CoT traces to produce a saliency reward that, when combined with outcome rewards in GRPO, trains LLMs to generate reasoning that genuinely influences final predictions.
CodeQ aggregates token rationales into code categories to enable global interpretability of LLMs, claiming over 50% entropy reduction and revealing model preference for syntactic cues plus human misalignment in a 37-person study.
Introduces BetXplain, an explanation-annotated dataset of social media betting ads collected from Instagram and Reddit for detecting manipulative and deceptive advertising.
SGR framework generates query-relevant subgraphs from knowledge graphs via schema-guided retrieval to guide LLM stepwise reasoning, reporting accuracy gains on QA benchmarks.
SGR enhances LLM reasoning accuracy by generating external subgraphs from knowledge bases and guiding progressive inference over them, yielding consistent gains over baselines on benchmarks.
citing papers explorer
-
Towards Faithful Agentic XAI: A Verification Method and an Open-World Benchmark for Better Model Faithfulness
FAX uses explicit claim verification against faithful tools to raise simulation faithfulness from 0.20 to 0.46 on the new CRAFTER-XAI-Bench open-world RL benchmark.
-
Geometry-Adaptive Explainer for Faithful Dictionary-Based Interpretability under Distribution Shift
GAE reduces the faithfulness gap in dictionary-based explainers under distribution shift by geometrically realigning the ID dictionary to the OOD-active subspace, with a quadratic excess-loss bound.
-
Beyond Accuracy: Robustness, Interpretability and Expressiveness of EEG Foundation Models
EEG foundation models show no single winner across failure modes, attend to correct brain regions but decode corrupted signals, and retain task information in early layers while late layers adapt during fine-tuning.
-
Monitoring Reasoning Models for Misbehavior and the Risks of Promoting Obfuscation
Chain-of-thought monitoring detects reward hacking in frontier reasoning models, but strong optimization against the monitor produces obfuscated misbehavior that remains hard to detect.
-
Evaluating Multi-turn Human-AI Interaction
Introduces the TCR framework to evaluate educational LLM assistants on transparency, consistency, and refinement in multi-turn interactions, complementing aggregate metrics.
-
Correcting Influence: Unboxing LLM Outputs with Orthogonal Latent Spaces
A latent mediation framework with sparse autoencoders enables non-additive token-level influence attribution in LLMs by learning orthogonal features and back-propagating attributions.
-
AtManRL: Towards Faithful Reasoning via Differentiable Attention Saliency
AtManRL learns an additive attention mask on CoT traces to produce a saliency reward that, when combined with outcome rewards in GRPO, trains LLMs to generate reasoning that genuinely influences final predictions.
-
Enabling Global, Human-Centered Explanations for LLMs:From Tokens to Interpretable Code and Test Generation
CodeQ aggregates token rationales into code categories to enable global interpretability of LLMs, claiming over 50% entropy reduction and revealing model preference for syntactic cues plus human misalignment in a 37-person study.
-
BetXplain: An Explanation-Annotated Dataset for Detecting Manipulative Betting Advertisements on Social Media
Introduces BetXplain, an explanation-annotated dataset of social media betting ads collected from Instagram and Reddit for detecting manipulative and deceptive advertising.
-
Stepwise Reasoning Enhancement for LLMs via External Subgraph Generation
SGR framework generates query-relevant subgraphs from knowledge graphs via schema-guided retrieval to guide LLM stepwise reasoning, reporting accuracy gains on QA benchmarks.
-
SGR: A Stepwise Reasoning Framework for LLMs with External Subgraph Generation
SGR enhances LLM reasoning accuracy by generating external subgraphs from knowledge bases and guiding progressive inference over them, yielding consistent gains over baselines on benchmarks.