{"id":"28ce16a7-5ac5-4490-b1d0-d9b2a8f70ad4","arxiv_id":"2410.03296","paper_version":4,"verdict":"UNVERDICTED","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"unknown","formal_verification":"none","parameter_count":0,"one_line_summary":"Self-explanations from LLMs produce faithful token subsets for correct predictions but align with human rationales only conditionally on text length and task complexity, unlike post-hoc attribution methods that highlight structural tokens.","lead":"The paper compares self-explanations generated by four open-weight LLMs against newly collected human rationales across sentiment classification, forced labour detection, and claim verification tasks, including multilingual variants. A smart generalist might read it to understand whether LLM-generated explanations are actually useful or faithful to how humans reason about text.","discovery_kind":"extension","skeptic_critique":{"model":"grok-4.3","headline":"Unvalidated Climate-Fever human rationale annotations as plausibility ground truth","rationale":"The reader's weakest_assumption directly identifies the load-bearing point for the plausibility component of the strongest_claim. Faithfulness evaluation (model-based) could be assessed separately, but the paper's headline distinction between explanation strategies requires both legs. Full-text access would allow checking the exact annotation section, but the abstract-level description already flags this as the least-secured premise.","tokens_in":1715,"tokens_out":357,"duration_ms":18027,"concrete_test":"Extract the Climate-Fever annotation protocol and raw per-annotator rationales from the paper; compute Cohen's kappa (or equivalent) across annotators. If kappa < 0.6, subsample 100 instances, re-annotate with two new annotators under the same guidelines, and recompute self-explanation alignment; a >15% shift in reported alignment metrics would indicate the plausibility results are unstable.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The central claim that self-explanations produce faithful token subsets (distinct from post-hoc methods) rests on two evaluations: (1) faithfulness w.r.t. correct model predictions and (2) plausibility via alignment with human rationales. The second leg depends on the newly collected Climate-Fever annotations. Without reported inter-annotator agreement, explicit annotation guidelines, number of annotators per instance, or bias-mitigation steps (e.g., blinding to model output), it is unclear whether these annotations constitute a stable, unbiased measure of human plausibility. If agreement is low or annotators default to surface cues, the reported dependence on text length/task complexity and the contrast with post-hoc methods lose support.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.3","summary":"The manuscript conducts a systematic empirical comparison of extractive self-explanations generated by four open-weight instruction-tuned LLMs against newly collected human rationales and post-hoc attribution methods across three text classification tasks (sentiment classification in English/Danish/Italian, forced labour detection, and claim verification on Climate-Fever). It evaluates plausibility through alignment with human annotations and faithfulness with respect to correct model predictions, concluding that self-explanation alignment with humans depends on text length and task complexity while self-explanations produce faithful token subsets, in contrast to post-hoc methods that emphasize structural and formatting tokens.","tokens_in":1822,"tokens_out":459,"duration_ms":12743,"significance":"If the human annotations prove reliable, the study offers useful evidence distinguishing LLM self-explanation strategies from post-hoc attribution in terms of faithfulness and plausibility. Strengths include the controlled multi-task and multi-language design, direct faithfulness checks against model predictions, and the collection of new human rationales for Climate-Fever to enable the plausibility comparison.","major_comments":[{"comment":"The section describing the annotation collection process for Climate-Fever: the plausibility evaluation and claims about dependence on text length/task complexity rest on these new human rationales, yet the manuscript reports neither inter-annotator agreement, the number of annotators per instance, explicit annotation guidelines, nor bias-mitigation procedures such as blinding annotators to model outputs. Without these, it is unclear whether the annotations constitute a stable ground truth for human plausibility.","section":"Annotation collection for Climate-Fever"}],"minor_comments":[{"comment":"The abstract and introduction could more explicitly state the number of models, tasks, and languages evaluated to improve scannability.","section":"Abstract"},{"comment":"Table or figure captions for the faithfulness and plausibility results should include the exact metrics used (e.g., token overlap, sufficiency) for immediate clarity.","section":"Results"}],"recommendation":"major_revision","confidential_remarks":"The manuscript is a reasonable fit for a computational linguistics venue focused on interpretability, but the unvalidated annotations are the primary load-bearing concern for the central plausibility claims."},"author_rebuttal":{"model":"grok-4.3","summary":"We thank the referee for their detailed review and constructive feedback on our manuscript. We address the major comment below and will revise the manuscript accordingly to improve the description of our human annotation process.","responses":[{"response":"We agree that additional details on the annotation collection process are required for the plausibility evaluation to be fully interpretable. In the revised manuscript, we will expand the relevant section to include inter-annotator agreement statistics, the number of annotators per instance, the full annotation guidelines, and bias-mitigation steps such as blinding procedures. These additions will clarify the stability of the human rationales as ground truth.","revision_made":"yes","referee_comment":"[Annotation collection for Climate-Fever] The section describing the annotation collection process for Climate-Fever: the plausibility evaluation and claims about dependence on text length/task complexity rest on these new human rationales, yet the manuscript reports neither inter-annotator agreement, the number of annotators per instance, explicit annotation guidelines, nor bias-mitigation procedures such as blinding annotators to model outputs. Without these, it is unclear whether the annotations constitute a stable ground truth for human plausibility."}],"tokens_in":1326,"tokens_out":262,"duration_ms":12684,"standing_objections":[]},"desk_editor":{"model":"grok-4.3","letter":"The core contribution here is a controlled empirical comparison of LLM self-explanations against human rationales on sentiment classification, forced labour detection, and claim verification. It includes Danish and Italian translations for one task and introduces fresh human annotations on Climate-Fever, then contrasts both with post-hoc attribution methods while checking faithfulness to correct model predictions. The finding that self-explanations select faithful token subsets while post-hoc methods favor structural tokens is a clear, usable distinction across the four open-weight models tested.","headline":"The paper adds new human rationale annotations for Climate-Fever and a multilingual comparison, but the plausibility evaluation rests on under-documented annotations.","tokens_in":2315,"tokens_out":169,"would_cite":false,"duration_ms":12323,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":{"model":"grok-4.3","evidence":[],"headline":"NLP/XAI evaluation of LLM self-explanations has no overlap with RS distinction-forcing or J-cost machinery","alignment":"orthogonal","rationale":"Paper studies plausibility/faithfulness of token rationales in text classification (SST, RaFoLa) via human annotations and LRP post-hoc methods. Central claims concern agreement metrics (Cohen's Kappa), POS/entity distributions, and masking-based faithfulness. No reference to or structural use of J(x) = ½(x + x⁻¹) − 1, phi-ladder, 8-tick periodicity, or the reality_from_one_distinction theorem chain. Domain is computational linguistics; RS has no theorems on rationale extraction or human-model agreement.","tokens_in":53791,"confidence":"high","tokens_out":164,"duration_ms":5782,"cache_read_input_tokens":38528,"cache_creation_input_tokens":0},"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.3","headline":"Self-explanations from LLMs yield faithful token subsets aligned with human rationales in text classification, unlike post-hoc methods that emphasize structural tokens.","keywords":["self-explanations","human rationales","text classification","LLM explanations","faithfulness","plausibility","post-hoc attributions","Climate-Fever"],"falsifier":"A controlled re-annotation of Climate-Fever examples by a second independent group of annotators that reverses the observed alignment ranking between self-explanations and post-hoc methods would falsify the claim of distinct explanation strategies.","tokens_in":2584,"feed_emoji":"🔍","tokens_out":701,"duration_ms":29339,"temperature":0.7,"pith_summary":"The paper tests whether instruction-tuned LLMs can produce usable explanations for their own text classification outputs by generating extractive self-explanations as input rationales. It compares these rationales to new and existing human annotations across sentiment classification in three languages, forced labour detection, and claim verification, with fresh human labels collected for the Climate-Fever dataset. Alignment between self-explanations and humans turns out to vary with input length and task difficulty, yet the self-explanations remain faithful to the tokens that actually drive correct model predictions. Post-hoc attribution techniques, by contrast, consistently surface formatting and structural tokens instead. The comparison therefore isolates two distinct explanation strategies that cannot be treated as interchangeable.","feed_headline":"LLM self-explanations match human rationales more faithfully than post-hoc methods","feed_subtitle":"Alignment varies with length and complexity across three tasks, but self-explanations stay tied to model outputs while post-hoc methods pick","key_machinery":"Extractive self-explanations (token rationales generated directly by the LLM) evaluated for both human plausibility and faithfulness to model predictions, contrasted with post-hoc attribution methods on the same inputs.","core_discovery":"Across sentiment, forced labour, and claim verification tasks, self-explanations generated by four open-weight LLMs produce token subsets whose faithfulness to correct model predictions exceeds that of post-hoc attributions; human alignment of these self-explanations depends on text length and task complexity, while post-hoc methods preferentially highlight structural and formatting tokens, indicating fundamentally different explanation strategies.","pith_inferences":["Applications that need explanations faithful to a model's actual decision process may benefit from using self-explanations rather than post-hoc attributions.","Hybrid systems could combine self-explanations for faithfulness with post-hoc methods for coverage of structural cues.","Standardized protocols for collecting human rationales would strengthen future comparisons of this kind.","The length- and complexity-dependence suggests testing self-explanation quality on longer documents or multi-sentence reasoning tasks next."],"forward_implications":["Alignment of self-explanations with human rationales is not uniform but scales with input length and task complexity.","Self-explanations remain faithful to the tokens supporting correct model outputs even when human agreement is only partial.","Post-hoc attribution methods systematically surface formatting and structural tokens rather than content tokens.","The pattern holds across English, Danish, and Italian versions of the sentiment task."],"fun_headline_variants":["LLM self-explanations yield more faithful rationales than post-hoc methods","Self-explanations align with human rationales depending on task and length","Post-hoc attributions favor structural tokens unlike LLM self-explanations","Faithfulness of self-explanations surpasses post-hoc across three tasks"],"cache_read_input_tokens":2112,"weakest_assumption_plain":"The newly collected human rationale annotations for Climate-Fever form a reliable and unbiased measure of plausibility for comparing LLM self-explanations.","fun_headline_variants_meta":{"raw":{"variants":["LLM self-explanations yield more faithful rationales than post-hoc methods","Self-explanations align with human rationales depending on task and length","Post-hoc attributions favor structural tokens unlike LLM self-explanations","Faithfulness of self-explanations surpasses post-hoc across three tasks"]},"model":"grok-4.3","cost_usd":0.008454,"raw_usage":{"total_tokens":3814,"prompt_tokens":650,"num_sources_used":0,"completion_tokens":71,"cost_in_usd_ticks":84537000,"prompt_tokens_details":{"text_tokens":650,"audio_tokens":0,"image_tokens":0,"cached_tokens":256},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":3093,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":650,"tokens_out":71,"duration_ms":25051,"temperature":1.0,"reasoning_tokens":3093,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-05-23T19:46:22.551538+00:00","model_set":{"reader":"grok-4.3"},"falsifier":"A controlled re-annotation of Climate-Fever examples by a second independent group of annotators that reverses the observed alignment ranking between self-explanations and post-hoc methods would falsify the claim of distinct explanation strategies.","supporting_citations":[],"review_version":1}