{"id":"b03d0114-bb31-445a-b910-85bef0d03881","arxiv_id":"2604.25866","paper_version":2,"verdict":"UNVERDICTED","confidence":"LOW","novelty_score":6.0,"correctness_risk":"unknown","formal_verification":"none","parameter_count":0,"one_line_summary":"Sparse autoencoders uncover that LLMs struggle with certain emotions because their causal features are more distributed, weaker, and co-activated with other emotions, with interventions improving recognition.","lead":"This paper applies sparse autoencoders to identify causal features inside LLMs that drive inference of specific emotions, showing that emotions like disgust rely on weaker, more distributed and overlapping features than concentrated ones like surprise or fear. The differences offer a mechanistic account for uneven model performance and are tested via targeted feature steering interventions.","discovery_kind":"new_application","skeptic_critique":{"model":"grok-4.3","headline":"Central claim asserted without any methods, data, or results in available text","rationale":"Reader correctly flags the abstract-only limitation and the unverified causality/intervention assumption as the weakest point; no independent evidence exists in the supplied text to test that assumption, so the verdict and low confidence are unchanged.","tokens_in":1694,"tokens_out":270,"duration_ms":15703,"concrete_test":"Obtain the full paper; inspect §3–4 for SAE training details, the exact intervention procedure (e.g., activation patching or steering vectors), and all result tables/figures reporting pre/post-intervention accuracy or feature statistics; if those sections are absent or show only correlational evidence, the causal claim remains unsupported.","verdict_should_be":"UNVERDICTED","load_bearing_attack":"The abstract states that SAEs identify causal sparse features with differing organizations across emotions (concentrated for surprise/fear, distributed/weaker/co-activating for disgust) and that two intervention experiments mitigate failures, but supplies zero details on SAE training, feature selection criteria, intervention protocol, datasets, baselines, or quantitative outcomes. The load-bearing step—establishing that identified features are causal drivers rather than correlates, and that steering them produces the claimed mitigation—therefore cannot be evaluated at all from the provided material.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.3","summary":"The manuscript claims that sparse autoencoders can be used to identify causal sparse features driving emotion inference in LLMs. It reports that emotions such as surprise and fear rely on highly concentrated feature sets, whereas disgust exhibits a distributed sparse causal organization in which its features are weaker, frequently co-activate with other emotions, and are overshadowed by anger features; these differences are presented as a mechanistic explanation for uneven model performance. The work further states that two intervention experiments—targeted steering of weaker causal features and global optimization of a steering vector over the identified features—mitigate emotion-specific failures and improve overall emotion recognition.","tokens_in":1787,"tokens_out":325,"duration_ms":23755,"significance":"If the reported causal organizations and intervention results hold, the work would supply a mechanistic account, grounded in sparse feature analysis, for why LLMs exhibit systematic biases across emotion categories and would demonstrate concrete steering-based remedies. Such findings could inform both interpretability research on affective representations and practical improvements in emotionally sensitive LLM applications.","major_comments":[{"comment":"Abstract: the central claims—that SAEs recover genuinely causal emotion features whose differing organizations explain performance gaps, and that two intervention experiments produce the stated mitigation and improvement—are asserted without any description of SAE training, feature selection criteria, intervention protocols, datasets, baselines, statistical tests, or quantitative outcomes. Because these elements are load-bearing for every substantive assertion, the manuscript supplies no basis on which the claims can be evaluated.","section":"Abstract"}],"minor_comments":[],"recommendation":"reject","confidential_remarks":null},"author_rebuttal":{"model":"grok-4.3","summary":"We thank the referee for their review. The sole major comment concerns the absence of methodological and experimental details in the abstract. We respond point by point below.","responses":[{"response":"The referee correctly observes that the provided manuscript consists only of the abstract, which contains no descriptions of SAE training, feature selection criteria, intervention protocols, datasets, baselines, statistical tests, or quantitative outcomes. Consequently, the central claims cannot be evaluated from the given text, and we have no additional manuscript sections to reference.","revision_made":"no","referee_comment":"[Abstract] Abstract: the central claims—that SAEs recover genuinely causal emotion features whose differing organizations explain performance gaps, and that two intervention experiments produce the stated mitigation and improvement—are asserted without any description of SAE training, feature selection criteria, intervention protocols, datasets, baselines, statistical tests, or quantitative outcomes. Because these elements are load-bearing for every substantive assertion, the manuscript supplies no basis on which the claims can be evaluated."}],"tokens_in":1298,"tokens_out":270,"duration_ms":34304,"standing_objections":["The full methodological and experimental details (SAE training, feature selection, intervention protocols, datasets, baselines, statistical tests, and quantitative outcomes) are not available because only the abstract is provided in the manuscript text."]},"desk_editor":{"model":"grok-4.3","letter":"The core issue is that this work exists only as an abstract. It claims SAEs reveal concentrated causal features for surprise and fear versus distributed, weaker, co-activating ones for disgust, plus two intervention experiments that fix the gaps. None of that can be checked because the text gives zero information on SAE training, feature selection, datasets, intervention protocols, baselines, or any numbers.\n\nWhat the abstract does sketch is a mechanistic angle on uneven emotion performance that prior work apparently missed. If the full paper delivered reproducible feature maps and intervention results that hold up, that would be a concrete step beyond correlation-style analyses.\n\nRight now the soft spot is total: the central claims rest on unshown evidence. The assumption that SAE features are causal drivers rather than correlates is stated but not demonstrated here, and the intervention results are mentioned without any protocol or outcome details. That makes the mechanistic explanation and the mitigation claims unevaluable.\n\nThis is for readers already working on SAE interpretability in LLMs who want to see whether the emotion-specific sparsity pattern holds once the methods appear. Until the full paper with code, data, and stats is available, it does not yet merit referee time.","headline":"Abstract-only paper asserts SAE-based causal differences across emotions but supplies no methods, data, or results to evaluate the claims.","tokens_in":2258,"tokens_out":310,"would_cite":false,"duration_ms":12456,"reading_group":"no","serious_thinker":"unclear","would_accept_peer_review":false},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.3","headline":"LLMs struggle with disgust because its causal features are weaker and more distributed than the concentrated sets used for surprise and fear.","keywords":["sparse autoencoders","causal features","emotion recognition","LLMs","mechanistic interpretability","disgust","surprise","fear"],"falsifier":"An intervention that activates or suppresses the identified disgust features produces no measurable change in the model's disgust classification accuracy on held-out examples.","tokens_in":2603,"feed_emoji":"🧠","tokens_out":584,"duration_ms":22467,"temperature":0.7,"pith_summary":"The paper applies sparse autoencoders to locate the specific features inside LLMs that causally drive emotion classification. It finds that surprise and fear each draw on tight clusters of strong features, while disgust draws on a scattered collection of weaker features that frequently overlap with those for anger. This difference in how the features are organized supplies a direct mechanistic account for the uneven accuracy across emotion categories. The authors then test two forms of intervention on these features to correct the identified weaknesses.","feed_headline":"Disgust uses weaker distributed features than surprise in LLMs","feed_subtitle":"Sparse autoencoders locate concentrated causal features for some emotions and scattered weaker ones for disgust, explaining accuracy gaps.","key_machinery":"causal sparse emotion features identified and organized by sparse autoencoders","core_discovery":"The central claim is that emotion inference in LLMs rests on sparse causal features whose organization varies by emotion: surprise and fear depend on highly concentrated feature sets, whereas disgust depends on a distributed organization in which its causal features are weaker, co-activate with features for other emotions, and are frequently overshadowed by causal features for anger; these representational differences explain the models' selective failures on certain emotions.","pith_inferences":["If feature concentration predicts accuracy, then other classification domains with uneven performance may also show concentrated versus distributed causal structures.","Disentangling co-activated features could become a general technique for improving reliability on any task where categories share representational overlap.","The intervention methods may transfer to non-emotion tasks once analogous causal features are located."],"forward_implications":["Targeted steering of the weaker causal features for disgust can reduce emotion-specific misclassifications.","Global optimization of a steering vector over all identified causal features raises overall emotion recognition accuracy.","Emotions whose features are concentrated are less susceptible to interference from other emotion features.","The same SAE-based analysis can be repeated on additional models to locate similar organizational differences."],"fun_headline_variants":["Disgust scatters weak features while surprise concentrates them in LLMs","Surprise and fear use concentrated features disgust does not in LLMs","Disgust causal features are weaker and coactivate with anger in LLMs","Distributed organization weakens disgust emotion inference in LLMs"],"cache_read_input_tokens":2112,"weakest_assumption_plain":"The features extracted by sparse autoencoders are genuine causal drivers of the model's emotion outputs rather than merely correlated patterns.","fun_headline_variants_meta":{"raw":{"variants":["Disgust scatters weak features while surprise concentrates them in LLMs","Surprise and fear use concentrated features disgust does not in LLMs","Disgust causal features are weaker and coactivate with anger in LLMs","Distributed organization weakens disgust emotion inference in LLMs"]},"model":"grok-4.3","cost_usd":0.009899,"raw_usage":{"total_tokens":4403,"prompt_tokens":672,"num_sources_used":0,"completion_tokens":64,"cost_in_usd_ticks":98987000,"prompt_tokens_details":{"text_tokens":672,"audio_tokens":0,"image_tokens":0,"cached_tokens":256},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":3667,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":672,"tokens_out":64,"duration_ms":39595,"temperature":1.0,"reasoning_tokens":3667,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-01T08:40:37.196642+00:00","model_set":{"reader":"grok-4.3"},"falsifier":"An intervention that activates or suppresses the identified disgust features produces no measurable change in the model's disgust classification accuracy on held-out examples.","supporting_citations":[],"review_version":2}