{"id":"fde8d2e5-daf7-4773-a468-7d2df6e48af3","arxiv_id":"2605.14694","paper_version":1,"verdict":"UNVERDICTED","confidence":"LOW","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"SAEs exhibit a rate-distortion-polysemanticity tradeoff where monosemanticity increases rate and distortion, with optimal polysemanticity set by feature co-occurrence probabilities in the data.","lead":"The paper shows that Sparse Autoencoders face an inherent tradeoff where forcing monosemantic features raises both the number of features used and reconstruction error. This matters for mechanistic interpretability because it frames polysemanticity as a data-distribution issue rather than purely an architectural one.","discovery_kind":"extension","skeptic_critique":{"model":"grok-4.3","headline":"The 'necessarily' claim for rate/distortion increase under monosemantic restriction holds only under toy generative assumptions, not shown generally.","rationale":"The reader's weakest_assumption directly identifies the same load-bearing point: the entire tradeoff demonstration rests on toy assumptions plus an explicit generative model. No stronger internal inconsistency is visible from the abstract and stated claims; the paper qualifies its theoretical results accordingly. The concrete_test above would falsify or confirm whether the necessity is an artifact of those assumptions.","tokens_in":1703,"tokens_out":311,"duration_ms":15361,"concrete_test":"Train SAEs on a controlled synthetic dataset whose feature co-occurrence probabilities are varied while keeping marginal feature probabilities fixed; compute the minimal rate and distortion achievable by an SAE constrained to be exactly monosemantic versus an unconstrained SAE. If the gap disappears for some co-occurrence regimes, the necessity does not follow from the generative model alone.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim states that restricting the SAE to be monosemantic necessarily increases rate and distortion. This is shown theoretically and empirically only under toy-modeling assumptions with an explicit generative model whose features have defined co-occurrence probabilities. The paper derives that optimal polysemanticity is set by those probabilities, but provides no argument that the necessity survives when the data-generating process is unknown or when features lack a clean generative structure. The real-world extension derives only necessary conditions on a polysemanticity measure, without demonstrating that the rate-distortion penalty must still occur.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.3","summary":"The paper introduces the Rate-Distortion-Polysemanticity tradeoff for Sparse Autoencoders. Under toy-modeling assumptions with an explicit generative model, it theoretically and empirically shows that restricting SAEs to monosemantic representations necessarily increases rate and distortion. It further shows that the polysemanticity of optimal SAEs is determined by the training data distribution, particularly feature co-occurrence probabilities. The work extends the analysis to real data by deriving necessary conditions that any polysemanticity measure must satisfy when the generative process is unknown, and benchmarks existing proxy metrics on SAEs trained on LLMs.","tokens_in":1847,"tokens_out":464,"duration_ms":19250,"significance":"If the results hold, the framing of polysemanticity as a data-distribution phenomenon (via co-occurrence) rather than purely an architectural failure would be a useful conceptual contribution to mechanistic interpretability. The derivation of necessary conditions for polysemanticity measures provides a concrete criterion that future metrics can be checked against. The paper receives credit for explicitly stating its toy-model assumptions and for attempting a bridge from the generative-model analysis to practical LLM SAEs.","major_comments":[{"comment":"Abstract and theoretical analysis: the central claim that monosemantic restriction 'necessarily' increases rate and distortion is derived only under toy generative assumptions with known feature co-occurrence probabilities. No argument is supplied showing that the necessity survives when the data-generating process is unknown (the realistic case), which is load-bearing for the claimed tradeoff applying to deployed SAEs.","section":"Abstract and theoretical analysis"},{"comment":"Empirical extension section: the benchmarking of proxy metrics on real LLM SAEs is presented at high level without reported error bars, baseline comparisons, or quantitative tables, so it is impossible to evaluate whether the necessary conditions are satisfied or whether the proxies behave as predicted by the toy analysis.","section":"Empirical extension section"}],"minor_comments":[{"comment":"Clarify the precise definitions of rate, distortion, and the polysemanticity measure used in the toy experiments so that the reported increases can be reproduced from the stated generative model.","section":"Toy-model experiments"}],"recommendation":"major_revision","confidential_remarks":null},"author_rebuttal":{"model":"grok-4.3","summary":"We thank the referee for their constructive feedback and for acknowledging the paper's contributions in framing polysemanticity as a data-distribution phenomenon and providing necessary conditions for polysemanticity measures. We address the major comments point by point below.","responses":[{"response":"We agree that the necessity claim for increased rate and distortion under monosemantic restriction is established only under the specified toy generative model assumptions, including known feature co-occurrence probabilities. The manuscript does not supply an argument demonstrating that this necessity holds in the general case where the data-generating process is unknown. The extension section instead focuses on deriving necessary conditions for any polysemanticity measure to be valid when the generative process is unknown. We will revise the manuscript to more explicitly state the scope of the necessity result and clarify that it does not directly apply to deployed SAEs without the toy assumptions.","revision_made":"partial","referee_comment":"[Abstract and theoretical analysis] Abstract and theoretical analysis: the central claim that monosemantic restriction 'necessarily' increases rate and distortion is derived only under toy generative assumptions with known feature co-occurrence probabilities. No argument is supplied showing that the necessity survives when the data-generating process is unknown (the realistic case), which is load-bearing for the claimed tradeoff applying to deployed SAEs."},{"response":"We acknowledge that the presentation of the benchmarking results for proxy metrics on SAEs trained on LLMs is at a high level and lacks error bars, baseline comparisons, and quantitative tables. This limits the ability to assess whether the necessary conditions are met or how the proxies align with the toy model predictions. In the revised version, we will expand this section to include error bars, relevant baseline comparisons, detailed quantitative tables, and explicit checks against the necessary conditions derived in the paper.","revision_made":"yes","referee_comment":"[Empirical extension section] Empirical extension section: the benchmarking of proxy metrics on real LLM SAEs is presented at high level without reported error bars, baseline comparisons, or quantitative tables, so it is impossible to evaluate whether the necessary conditions are satisfied or whether the proxies behave as predicted by the toy analysis."}],"tokens_in":1389,"tokens_out":457,"duration_ms":32514,"standing_objections":[]},"desk_editor":{"model":"grok-4.3","letter":"The core contribution is the explicit Rate-Distortion-Polysemanticity tradeoff. Under the paper's toy generative model, they show that forcing an SAE to be monosemantic raises both rate and distortion, and that the optimal amount of polysemanticity is set by the probability that features co-occur in the data. The derivation is straightforward and the toy experiments back it up without obvious gaps.\n\nThat framing is new relative to earlier SAE work and gives a data-dependent explanation for why polysemanticity appears. It shifts the question from pure architecture fixes toward whether the training distribution itself forces some overlap.\n\nThe real-data part is weaker. They derive necessary conditions a polysemanticity measure must satisfy when the generative process is unknown, then benchmark a few proxy metrics on LLMs. This is useful as a starting point but does not demonstrate that the rate-distortion penalty still holds outside the toy setting. No error bars or strong baselines are described, and the extension stays at the level of conditions rather than a direct test.\n\nThe central claims are carefully scoped to the toy assumptions, so there is no overclaim on that front. The math and citation pattern look standard for the area.\n\nThis is worth a reading group for anyone working on SAEs or mechanistic interpretability who wants to think about data limits. It deserves peer review because the toy result is clean and the question it raises is concrete, even if the LLM experiments need more development to carry the same weight.","headline":"The paper cleanly derives a rate-distortion cost to monosemantic SAEs under toy generative models and ties optimal polysemanticity to feature co-occurrence, but the real-data section only gives necessary conditions without showing the tradeoff persists.","tokens_in":2305,"tokens_out":387,"would_cite":false,"duration_ms":15007,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.3","headline":"Sparse autoencoders must trade higher rate and distortion for monosemantic features.","keywords":["sparse autoencoders","polysemanticity","rate-distortion tradeoff","mechanistic interpretability","large language models","generative models","feature co-occurrence"],"falsifier":"Construct data in which every pair of features has zero co-occurrence probability, train both monosemantic and unrestricted SAEs, and check whether the monosemantic version can match the rate and distortion of the unrestricted version.","tokens_in":2621,"feed_emoji":"⚖️","tokens_out":441,"duration_ms":20517,"temperature":0.7,"pith_summary":"The paper establishes that sparse autoencoders cannot simultaneously minimize rate, minimize distortion, and enforce monosemantic features. Restricting an SAE to monosemantic representations forces an increase in the number of active features and in reconstruction error. This tradeoff holds because the optimal degree of polysemanticity is fixed by feature co-occurrence probabilities under an assumed generative model of the inputs. A reader would care because the result reframes polysemanticity as a property of the data distribution rather than a pure failure of the SAE architecture or training procedure.","feed_headline":"Monosemantic SAEs raise rate and distortion","feed_subtitle":"Optimal polysemanticity is fixed by how often features co-occur in the training distribution","key_machinery":"The rate-distortion-polysemanticity tradeoff, which shows that monosemantic constraints on SAEs increase rate and distortion because optimal polysemanticity is set by feature co-occurrence probabilities in a generative model of the data.","core_discovery":"Under toy-modeling assumptions, restricting the SAE to be monosemantic necessarily comes with an increase in rate and distortion. Assuming a generative model behind the input observations, the degree of polysemanticity of optimal SAEs is determined by the training data distribution, especially by the probability of features to co-occur.","pith_inferences":[],"forward_implications":[],"fun_headline_variants":["Monosemantic SAEs increase rate and distortion","Polysemanticity in SAEs depends on feature co-occurrence","Training data sets optimal SAE polysemanticity","SAE tradeoff links monosemanticity to higher rate","Optimal SAEs polysemantic based on data distribution"],"cache_read_input_tokens":64,"weakest_assumption_plain":"The analysis depends on toy-modeling assumptions and the existence of an underlying generative model that sets polysemanticity through feature co-occurrence probabilities.","fun_headline_variants_meta":{"raw":{"variants":["Monosemantic SAEs increase rate and distortion","Polysemanticity in SAEs depends on feature co-occurrence","Training data sets optimal SAE polysemanticity","SAE tradeoff links monosemanticity to higher rate","Optimal SAEs polysemantic based on data distribution"]},"model":"grok-4.3","cost_usd":0.005139,"raw_usage":{"total_tokens":2400,"prompt_tokens":635,"num_sources_used":0,"completion_tokens":72,"cost_in_usd_ticks":51390500,"prompt_tokens_details":{"text_tokens":635,"audio_tokens":0,"image_tokens":0,"cached_tokens":64},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":1693,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":635,"tokens_out":72,"duration_ms":14655,"temperature":1.0,"reasoning_tokens":1693,"cache_read_input_tokens":64,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-06-30T21:10:53.958834+00:00","model_set":{"reader":"grok-4.3"},"falsifier":"Construct data in which every pair of features has zero co-occurrence probability, train both monosemantic and unrestricted SAEs, and check whether the monosemantic version can match the rate and distortion of the unrestricted version.","supporting_citations":[],"review_version":1}