{"id":"658667c0-8c63-4031-b4f5-b01f130d887a","arxiv_id":"2607.23379","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"Activation Oracles trained on Taboo subjects selectively fail to verbalize the concept present during their own training, even when that concept remains linearly decodable inside the oracle.","lead":"Fine-tuned Activation Oracles can become concept-specific anti-readers: they fail to report the hidden concept that was always present in their training data. This matters because learned interpretability tools may silently omit exactly the information auditors most want to recover.","discovery_kind":"extension","skeptic_critique":{"model":"moonshotai/kimi-k3","headline":"The behavioral anti-reading effect is well supported, but the \"readout-side suppression\" mechanism rests on AO-internal probes that may decode subject-checkpoint identity rather than a functionally extracted concept, and on ablation \"restoration\" that is relative and partial.","rationale":"The reader's weakest assumption concerned external validity (single backbone, synthetic Taboo organisms, LoRA-only, unshipped code, LLM-judge metrics). I agree those are the right reasons for CONDITIONAL rather than ACCEPT, but the most load-bearing internal soft spot is different: the mechanistic leg of the strongest claim. The reader's verdict and confidence already reflect the paper's hedged framing (\"indicate,\" \"suggest\"), and the paper states the generalization limit itself in §A, so I do not think the verdict should move in either direction. The behavioral core — concept-specific anti-reading with a multi-concept control that defeats the obvious checkpoint-identity confound — is strong enough that REJECT or a downgrade would be unwarranted. Conversely, ACCEPT would overstate the mechanistic evidence: the probe decodability result carries an unaddressed grouped-CV confound the authors themselves raise in a parallel context, and the ablation restoration is relative and partial. The proposed grouped cross-protocol probe test is cheap (it reuses existing captured activations and the existing probe pipeline) and would directly settle whether \"the concept remains inside the FT-AO\" is a functional-representation claim or a fingerprint claim; the absolute-P(c⋆) ablation comparison would calibrate how much weight the L18–23 localization can bear. Until those are run, CONDITIONAL with high confidence in the synthetic setting, as the reader assigned, is the right landing point.","tokens_in":37624,"tokens_out":5017,"duration_ms":208893,"concrete_test":"Rerun the AO-internal probe analysis (App. D) with checkpoint-grouped evaluation: for each (regime, AO, layer) cell, train the 5-way concept probe on AO hidden states from injections of cooperative subjects and test on strict subjects (and vice versa). If cross-protocol transfer stays far above the 20% chance baseline, the concept genuinely persists inside the FT-AO and the readout-side interpretation survives; if it collapses, §7.1's decodability claim reflects checkpoint fingerprints and must be weakened. Complementarily, report absolute P(c⋆) after L18–23 ablation alongside base-AO P(c⋆) on the same captures: if post-ablation P(c⋆) approaches base-AO levels, the localization claim is meaningfully supported; if it remains near the intact-FT-AO floor, \"restoration\" is a small relative effect.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The strongest claim has two legs: (1) own FT-AOs selectively fail to verbalize their training concept; (2) this is readout-side suppression, not absence of information — the concept remains decodable inside the FT-AO and suppression localizes to mid-layer LoRA updates (L18–23). Leg 1 is solid: the own/cross/base gap appears in distributional metrics (P(c⋆), rank; np≈300) that do not depend on the Sonnet judge, and the leaf_moon multi-concept control (§8) rules out checkpoint-identity explanations of the blind spot, since one checkpoint yields differential leaf-vs-moon reading. Leg 2 is softer. (a) The AO-internal probes (App. D, Table 2) pool hidden states from injections of 10 subject checkpoints and use ungrouped 5-fold CV — the exact confound the authors themselves flag for subject probes in §5.2 (\"examples from the same fine-tuned checkpoints occur across cross-validation folds\") but do not address for the AO-internal case. If the L33 probe classifies which checkpoint's activations were injected rather than a concept the AO functionally extracted, \"the target remains decodable inside the oracle\" holds only in a weak fingerprint sense, and the probe-vs-LogitLens contrast cannot distinguish readout suppression from failed extraction. (b) The ablations (App. H, Table 7) are blunt: zeroing lora_B over 6-layer ranges, with L0–5 ablation collapsing P(c⋆) to ~0 for every concept, showing the updates are globally entangled. The L18–23 \"restoration\" is ρ≈1.5–3.0 off already-low baselines (0.09–0.36), so absolute recovery stays far below base-AO levels (~0.6–0.75); \"localizes to L18–23\" is a relative statement, not evidence those layers implement suppression. Net: the central behavioral phenomenon stands; the mechanistic localization is suggestive, not established, and no code is shipped to verify it.","agreement_with_reader":"partial"},"referee_report":{"model":"moonshotai/kimi-k3","summary":"The paper studies Activation Oracles (AOs) — LLMs trained to answer natural-language questions about another model's residual-stream activations — in a controlled Taboo Word Guessing model organism. Subjects (Qwen3-8B LoRA adapters) are fine-tuned to internally use one of five hidden concepts while avoiding disclosure, under cooperative/strict protocols and two mixture strengths. The central finding is that an AO fine-tuned on a Taboo subject does not become a specialist reader of that subject's hidden concept; instead it becomes a concept-specific anti-reader, recovering its own training concept worse than a base AO and worse than cross FT-AOs trained on other concepts (diagonal suppression in the concept×concept heatmap, Fig. 6). The effect persists for multi-concept (leaf_moon) subjects (§8), ruling out a checkpoint-identity account of the blind spot. Mechanistically, the authors argue the failure is readout-side: the concept remains linearly decodable from AO hidden states (App. D), target rank degrades in late AO layers (Fig. 8), and zeroing LoRA updates in layers L12–23 partially restores target probability (Fig. 9, Table 7).","tokens_in":38098,"tokens_out":5081,"duration_ms":193028,"significance":"If the result holds, it is a useful and well-controlled cautionary result for the growing line of activation-to-language interfaces (AOs, NLAs): a learned reader can acquire concept-specific reporting blind spots induced by its own training distribution, so representation-level decodability, behavioral leakage, and AO-verbalizability must be evaluated separately. Strengths worth naming: the behavioral anti-reading claim is supported by distributional metrics (P(c⋆), rank, np≈300 per cell) that do not depend on the LLM judge; the design includes genuine negative controls (OFFTOPIC, SAMETEXT teacher-forced neutral text) and a multi-concept control that cleanly separates concept-specific from subject-specific failure; the own/cross/base three-way comparison is a non-tautological external contrast; and the appendices ship unusually complete reproducibility detail (prompt templates, seed schedules, LoRA hyperparameters, per-cell tables). The phenomenon is falsifiable and the claims are mostly hedged appropriately. Generality is limited by construction (single backbone, five concrete nouns, LoRA-only, synthetic training loop), which the authors state plainly in App. A.","major_comments":[{"comment":"The AO-internal probe evidence has the exact confound the authors themselves flag for subject-level probes in §5.2 and Appendix C ('the cross-validation split is not grouped by checkpoint... may reflect both concept-correlated structure and stable checkpoint-level signatures'), but the caveat is not carried over. The AO-internal probes pool hidden states from injections of only 10 subject checkpoints (2 per concept) with 30 prompts each, using ungrouped stratified 5-fold CV. Examples from the same checkpoint therefore appear in train and test folds, and with only two checkpoints per concept label, 'decodes the concept' and 'decodes which of two checkpoint distributions was injected' are nearly indistinguishable. This matters because §7.1 and D.2 draw a strong directional conclusion from these probes: 'anti-reading is not well explained as a failure to extract the hidden concept from the","section":"§7.1 / Appendix D.1–D.2, Table 2"},{"comment":"The ablation evidence is presented as causal localization of concept-specific suppression, but three gaps weaken it. (i) Only own-concept P(c⋆) is measured after ablation; there is no control showing that ablating L18–23 does not equally boost non-own concepts. Since ablation partially resets the FT-AO toward base behavior, and the base AO reads everything well, a generic partial reset predicts the observed restoration without any concept-specific mechanism. Reporting post-ablation probabilities for the other four concepts (or running the same ablation on cross FT-AO/subject pairs) would isolate the concept-specific component. (ii) The restoration is relative to already-suppressed baselines: e.g., leaf-FT baseline P(c⋆)=0.143 with peak ρ=2.01 gives absolute P(c⋆)≈0.29, still far below the base AO's 0.63 (Table 26). 'Restoration' should be reported in absolute terms alongside ρ. (iii) n=2","section":"§7.3 / Appendix H, Table 7, Fig. 9"},{"comment":"A substantial part of the own-FT-AO degradation is generic fine-tuning degradation rather than concept-specific anti-reading: for cooperative α=1.0, cross FT-AOs already drop from 94% (base) to 71% exact recovery, with own FT-AOs at 35%. The concept-specific effect is the own-vs-cross gap, not the base-vs-own gap, and the SAMETEXT regime drives much of it (Table 22: cross 0–35%, own 0–1%, while base retains 34–98%). The aggregate framing in §6.1 ('exact recovery drops from 94% for the base AO and 71% for cross FT-AOs to 35% for own FT-AOs') is accurate but the text would be clearer if it explicitly decomposed the total drop into a generic FT term and a concept-specific term, and stated which regimes carry the concept-specific component. This is a framing/quantification issue, not a correctness error, but it is load-bearing for how large the claimed phenomenon is.","section":"§6.1, Fig. 5"}],"minor_comments":[{"comment":"Typo: 'becomeconcept-specific anti-readers' (missing space).","section":"Abstract"},{"comment":"Caption is truncated mid-sentence: 'Median LogitLens rank of the target token when the AO's own Qwen3-8B' — the sentence is incomplete.","section":"Appendix D.3, Table 3 caption"},{"comment":"The SAMETEXT row labels 'moon-FT (own)' for an evaluation on the leaf subject; moon-FT is a cross AO for the leaf subject, so the '(own)' tag appears to be an error. Please check.","section":"Appendix G.3, Table 6"},{"comment":"SAMETEXT stores three identical copies per neutral text so that 'downstream aggregation uses the same nominal capture count.' These copies are not independent samples; please state explicitly how they are treated in the np≈300 counts and in the bootstrap CIs of Fig. 5, to rule out pseudoreplication.","section":"Appendix E.1 (SAMETEXT sampling)"},{"comment":"Substring exact match counts 'leafy' and 'books' as exact recoveries. This is mostly harmless given the Sonnet tier rubric, but worth one sentence noting the direction of the bias (it can only inflate exact recovery, including for own FT-AOs, so it works against the anti-reading claim — actually a point in the paper's favor).","section":"Appendix F.2"},{"comment":"Judge sample size nj=25 per cell gives 4% granularity on exact/semantic recovery; all judge-based percentages are multiples of 4. The distributional metrics (np≈300) mitigate this, but the resolution limit should be noted where judge numbers are quoted to the percent.","section":"Appendix F.3"},{"comment":"The subscript notation (e.g., '[48,51] 2') indicating fewer than four cross AOs is only defined in the J.4 preamble; several strict α=0.5 own cells (leaf, moon) are missing entirely (dashes). Please explain why those FT-AOs were unavailable and define the subscript at first use.","section":"Tables 23–24 (App. J.4)"},{"comment":"'activations from layers at roughly 25, 50 and 75 percents of the AO stack' — grammar; also worth stating explicitly that the fixed evaluation layer ℓ=18 of 36 corresponds to the 50% training injection point, so evaluation is in-distribution for the injection layer.","section":"Appendix B.3"},{"comment":"The caption describes panels (a) and (b) with left/right sub-panels, but the layout and what is plotted in (b) (which metric, which aggregation) are not specified; please make the figure self-contained.","section":"Figure 1"},{"comment":"'nostalgebraist (2020)' is cited with lowercase and as a blog post — acceptable, but consider consistent formatting; also double-check the Karvonen et al. arXiv identifier (abs/2512.15674).","section":"References"}],"recommendation":"major_revision","confidential_remarks":"The behavioral anti-reading phenomenon (own-vs-cross diagonal suppression, multi-concept control) looks solid and is, in my view, the paper's real contribution; I would not want the revision burden to exceed re-running existing pipelines. The two requested additions (checkpoint-grouped CV for the AO-internal probes; cross-concept controls and absolute values for the ablations) are cheap given the infrastructure the authors clearly have, but either could in principle weaken the 'readout-side suppression' framing of §7, which is why I land on major rather than minor revision. The synthetic single-backbone setting is acknowledged honestly and is appropriate for a model-organism paper; I have no concerns about the citation pattern or scope fit."},"author_rebuttal":null,"desk_editor":{"model":"grok-4.5","letter":"The thing worth knowing is the counterintuitive empirical result: when you train an Activation Oracle on a Taboo subject that hides concept c*, the oracle gets selectively worse at saying c*, not better. Base and cross FT-AOs still recover it; own FT-AOs do not. That own-vs-cross diagonal, plus the leaf/moon multi-concept check, is the real contribution. Prior AO work showed you can read fine-tuned secrets; this shows the reader can learn a concept-specific blind spot from the same training loop.\n\nThey run a clean chain for the behavioral claim. Behavioral non-leak, subject probes and Δ-LogitLens, base-AO recovery, then own/cross/base on exact recovery, P(c*), and rank. Distributional metrics with ~300 captures do not depend on the Sonnet judge. Multi-concept subjects kill the cheap “the AO just hates that checkpoint” story. Appendices are thick and the comparisons are the right ones.\n\nSoft spots, in proportion. The setting is synthetic (Qwen3-8B LoRA, five nouns, fixed inject layer). They say so. No code or data shipped, which matters for a result this dependent on training details. The mechanism leg is weaker than the abstract implies. AO-internal probes use ungrouped CV over injections from multiple checkpoints—the same confound they flag for subject probes—so “still decodable inside the oracle” may be fingerprinting rather than functional extraction. Ablations zero whole six-layer LoRA blocks; L0–5 kills everything, and L18–23 “restoration” is relative to already-low own baselines, not recovery to base-AO levels. So readout-side localization is suggestive, not established. Entropy analysis helps a bit (confident wrong token, not pure uncertainty).\n\nCitations look appropriate (Karvonen, Cywinski, ELK, LogitLens, causal tracing). Math is standard probe/ablation work, not load-bearing theory.\n\nWho it’s for: people building or auditing activation-to-language tools and ELK-style readers. Worth a serious referee. I’d bring it to reading group and cite the anti-reading phenomenon if I work on AOs; I’d treat the circuit story as a hypothesis until someone ships code or cleaner interventions.","headline":"Solid own-vs-cross anti-reading result in a Taboo AO setup; the behavioral finding is real, the readout-side story is only half-nailed.","tokens_in":38986,"tokens_out":570,"would_cite":true,"duration_ms":18044,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.5","headline":"Fine-tuned Activation Oracles can become concept-specific anti-readers that fail to report the hidden concept present throughout their own training.","keywords":["activation oracles","interpretability","latent knowledge","fine-tuning","readout failure","Taboo word guessing","concept-specific blind spots"],"falsifier":"Train own and cross Activation Oracles on the same Taboo subjects and check whether own-oracle recovery of the matched concept stays below base and cross oracles while linear probes still decode the concept inside the own oracle and ablating mid-layer LoRA updates restores target probability.","tokens_in":38680,"feed_emoji":"🔍","tokens_out":896,"duration_ms":23800,"temperature":0.7,"pith_summary":"Activation Oracles are language models trained to answer natural-language questions about another model’s internal activations, offered as a flexible way to read information that is present inside a model but missing from its text. This paper shows that those oracles are not neutral meters: when trained on a subject that hides a known concept while still representing it internally, the oracle can selectively lose the ability to name that same concept. In a controlled Taboo Word Guessing setup, own fine-tuned oracles underperform both a base oracle and oracles trained on other concepts, even though the target remains linearly decodable inside the oracle and recoverable by simpler readouts. The failure localizes to the oracle’s readout path rather than to missing information in the subject. The practical point is that behavioral secrecy, representation-level decodability, and oracle-verbalizability can come apart, so a learned interpretability interface can acquire blind spots induced by its training loop.","feed_headline":"Oracles trained to read hidden concepts can learn not to","feed_subtitle":"Fine-tuned activation readers develop concept-specific blind spots even when the target stays decodable inside them.","key_machinery":"Concept-specific anti-reading in Activation Oracles: LoRA-trained language models that answer questions about injected residual-stream activations, evaluated in own-versus-cross comparisons on Taboo subjects, with probes, LogitLens ranks, and mid-layer LoRA ablations separating representation decodability from verbalization.","core_discovery":"Contrary to the expectation that an Activation Oracle trained on a fine-tuned subject becomes a specialist reader of that subject’s hidden concept, fine-tuned oracles can become concept-specific anti-readers: they selectively fail to recover the concept that was persistently present during their own training, while often retaining recovery of other concepts. The target remains decodable from subject activations and inside the oracle itself; suppression arises in the oracle readout pathway.","pith_inferences":["Any learned monitor trained only on activations from a model that systematically conceals a feature may quietly learn not to name that feature.","Similar anti-reading could appear in other activation-to-text interfaces if training repeatedly pairs the same latent with non-disclosure answers.","Safety evaluations that rely on fine-tuned elicitors may need held-out concept suites and cross-trained controls before claiming coverage.","If mid-layer readout updates mediate the failure, interventions that freeze or diversify that path might reduce concept-specific omission."],"forward_implications":["A learned activation reader can acquire reporting policies that omit concepts present in both subject and oracle representations.","Checking that information is decodable in the subject is not enough to trust an oracle’s answers about it.","Own-versus-cross training comparisons become a necessary reliability test for activation-to-language interfaces.","Blind spots can follow the oracle’s training concept even when multiple hidden concepts sit in the same subject activations.","Auditing tools that are themselves fine-tuned may need readout-path diagnostics, not only subject-side probes."],"fun_headline_variants":["Fine-tuned oracles learn concept-specific blind spots","Activation oracles can become anti-readers of their training concept","Oracles fail to read the hidden concept they trained on","Readout pathway suppresses concepts still decodable inside oracles","Behavioral leakage and AO verbalizability come apart"],"cache_read_input_tokens":32896,"weakest_assumption_plain":"That a controlled Taboo Word Guessing setup on one backbone with a few hidden nouns is a fair enough stand-in for the hidden-information settings where activation oracles would be used as auditors.","fun_headline_variants_meta":{"raw":{"variants":["Fine-tuned oracles learn concept-specific blind spots","Activation oracles can become anti-readers of their training concept","Oracles fail to read the hidden concept they trained on","Readout pathway suppresses concepts still decodable inside oracles","Behavioral leakage and AO verbalizability come apart"]},"model":"grok-4.5","effort":"low","cost_usd":0.003426,"raw_usage":{"total_tokens":1132,"prompt_tokens":787,"num_sources_used":0,"completion_tokens":66,"cost_in_usd_ticks":34264000,"prompt_tokens_details":{"text_tokens":787,"audio_tokens":0,"image_tokens":0,"cached_tokens":128},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":279,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":787,"tokens_out":66,"duration_ms":6268,"temperature":1.0,"reasoning_tokens":279,"cache_read_input_tokens":128,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-30T23:49:05.513561+00:00","model_set":{"reader":"grok-4.5"},"falsifier":"Train own and cross Activation Oracles on the same Taboo subjects and check whether own-oracle recovery of the matched concept stays below base and cross oracles while linear probes still decode the concept inside the own oracle and ablating mid-layer LoRA updates restores target probability.","supporting_citations":[],"review_version":1}