{"id":"083fb84d-753a-4208-a554-6a0838ea4854","arxiv_id":"2607.14103","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"Latent communication between LLM agents preserves far more SAE features than text, but those extra features encode surface form and provide no task-level advantage on text-expressible tasks.","lead":"This paper tests whether AI agents communicating through internal model states instead of text can preserve more information, and finds that while latent states carry extra features, these features are not useful for the tasks tested. The result is an honest negative: for text-expressible tasks, standard text communication remains as good as direct latent-state communication.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The 88% feature-loss result compares contextual prompts to bare concept names, not to natural-language passages; the claim that text destroys only surface form is not supported for the text-passage channel.","rationale":"The reader's verdict correctly flags external validity: all tasks are text-expressible. I find a more immediate internal validity problem: the feature-loss statistic that anchors the negative conclusion is computed against bare concept names, while the paper's own task-level evaluation shows that realistic text communication (text-passage) is the strong baseline. The 88% number therefore conflates context removal with text serialization. The augmentation experiment, presented as decisive, uses the same bare-name-derived 'lost features' and lacks a random-feature control, so its degradation result does not uniquely support the surface-form interpretation. This reinforces the reader's conditional verdict: the paper is an honest, well-structured negative result for its specific tasks, but the headline 'limits of text' and the 'lost features are surface form' claim go beyond the evidence. A passage-based round-trip and a random-feature control would settle whether the central negative claim survives. Because the task-level results and alignment findings remain useful, I do not recommend rejection; the verdict stays conditional (UNCHANGED relative to the reader).","tokens_in":12953,"tokens_out":6236,"duration_ms":60825,"concrete_test":"Run the §2.5 feature-survival analysis with the text-passage condition (e.g., 'The process by which plants convert sunlight into glucose is called Photosynthesis') as the text-channel re-encoding instead of the bare name, and repeat the §3.6 augmentation with the features lost in that passage-based round-trip. If survival rate rises substantially above 11.7% or the lost-feature augmentation no longer degrades accuracy, the conclusion that text destroys only surface form is unsupported. Include a random-feature control matched in dimension and norm in the augmentation; if random features degrade equally, the current lost-feature result is an artifact of injection or alignment.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central negative claim—that text serialization destroys 88% of SAE features and that the lost features encode surface form rather than task-relevant semantics—rests on a text round-trip that compares a contextual prompt to the bare concept name (§2.5, Step 2). This is not text serialization as used in the task-level evaluation: the realistic text channel is text-passage, which transmits a full descriptive passage and outperforms the latent channels (§3.5). The 88% destruction rate therefore measures the effect of removing context, not the effect of converting a latent representation into natural language. The augmentation experiment (§3.6) inherits this mismatch: the 'lost features' are features active in contextual encoding but absent from bare-name encoding, and adding them to a text-passage representation degrades performance. That result shows those context-only features are not useful on top of a rich passage; it does not show they are surface form. A random-feature control is also missing, so the degradation could reflect Procrustes misalignment of out-of-distribution features rather than their semantic content. Consequently the strongest claim is not established for the text channel that actually matters.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper constructs three inter-agent communication channels—dense latent, SAE-sparse, and text—and compares them on concept-discrimination tasks involving Llama 3.1 8B-Instruct and Mistral 7B. The central positive results are that the SAE-sparse channel preserves high probe accuracy at substantial compression, and that Procrustes alignment enables cross-architecture concept retrieval well above chance. The central negative claim is that text serialization destroys 88% of SAE features, but that the lost features encode surface form rather than task-relevant semantics, so the latent channel never outperforms text on the tested tasks. The paper is structured as a gated experimental study with explicit prerequisites, claims C1–C3, task-level evaluation, and text-augmentation tests, and it is candid about the limitation that all tasks have text-expressible answers.","tokens_in":13240,"tokens_out":2258,"duration_ms":24470,"significance":"If the central claims hold, the paper makes a useful contribution by quantifying, at feature level, how much information is lost in text-based agent communication and by providing strong evidence for cross-architecture representational convergence (92% top-1 retrieval against a 0.87% chance baseline). The methodological apparatus—Procrustes with held-out generalization, SAE feature survival, text augmentation—is well conceived and goes beyond simple cosine-similarity comparisons. The honest negative result is valuable for the multi-agent systems community. However, the strongest negative conclusion depends on a feature-loss measurement that is currently performed on bare concept names rather than on the realistic text-passage channel; this needs to be fixed or the claims must be substantially weakened.","major_comments":[{"comment":"The 88% feature-loss rate is computed by comparing a contextual cloze-prompt encoding with a bare concept-name encoding (e.g., \"Photosynthesis\"). This measures the effect of removing context, not the effect of serializing the sender's latent representation into a natural-language passage. The task-level evaluation in §3.5 uses text-passage as the realistic text channel, and this channel outperforms the latent channels on all task types. Therefore the central claim that 'text serialization destroys 88% of SAE features' is not established for the text channel that actually matters. The authors should recompute the feature-survival analysis using the textual passages used in §2.7, or explicitly reframe the 88% result as 'context removal destroys features' and separate it from the serialization claim.","section":"§2.5, §3.3, Table 2"},{"comment":"The lost-feature augmentation experiment lacks a random-feature control. The lost features are defined as those active in contextual encoding but absent from bare-name encoding; when added to a text-passage representation, they degrade performance. Without a control of randomly selected SAE features of matched activation frequency and magnitude, the degradation cannot be attributed to the semantic content of the lost features. It may instead reflect that these features are out-of-distribution relative to the receiver's text-passage representation, or that the Procrustes mapping is poorly calibrated for them. A random-feature control is necessary to support the conclusion that the lost features 'act as noise' because they are surface-form features.","section":"§3.6, Eq. (2)"},{"comment":"The headline C1 result—99.4% SAE-sparse vs 80.4% text accuracy—is reported with a text-channel standard deviation of 24.4%. No significance test is provided, and with this variance the 19.4pp gap may not be robust across concepts or prompt samples. The authors should report confidence intervals or a paired test across the 84 concept pairs, and clarify how probe accuracy varies across concepts.","section":"§3.2, Table 1"}],"minor_comments":[{"comment":"The abstract states 'text serialization destroys 88% of SAE features' without the important qualifier that this is measured against bare-name re-encoding, not passage-level serialization. The limitations section does acknowledge the text-expressibility assumption, but the abstract and §3.3 overstate the specificity of the result.","section":"Abstract / §4.4"},{"comment":"The cross-lingual concept result (+1.4pp over text-passage) is reported as the only regime where latent matches text. Given the small number of tasks (74), it would be useful to report a confidence interval or significance test; without it, the 'parity' claim is not statistically grounded.","section":"§2.7 / §3.5"},{"comment":"The 92% top-1 retrieval at 140 anchors is based on only 25 held-out concepts. The authors should acknowledge the small holdout size and report the binomial confidence interval; at n=25, the interval around 92% is roughly [74%, 99%].","section":"§3.4"},{"comment":"Several typographical issues: 'LLMS' in §1, 'an corresponding' in §4.4, 'T able' in Table captions, 'identityreplacement' in §3.3. Figure 1 caption is incomplete ('LLama... offers different positions'). The paper should be copy-edited before publication.","section":"Throughout"}],"recommendation":"major_revision","confidential_remarks":"The paper has a solid gated design and the positive results on cross-architecture alignment and SAE channel fidelity are valuable. The main issue is that the negative conclusion—'text destroys mostly surface form'—rests on a text round-trip that is not the text channel used in the task-level comparisons. This is fixable with additional experiments, so I recommend major revision rather than rejection. I would also encourage the authors to add the random-feature control in the augmentation experiment, as it is directly load-bearing for the central claim."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Dear colleague,\n\nHere’s my take. The paper’s real contribution is the measurement framework: three channels (dense, SAE-sparse, text), a Procrustes alignment tested on held-out concepts, and a task-level protocol that avoids training a new probe at inference time. The 92% top-1 cross-architecture retrieval at 140 anchors, against a 0.87% random baseline, is a solid result. The task-level evaluation is also honestly done: on all tested tasks, text-passage matches or beats the latent channels, and the augmentation experiment shows no benefit from adding latent features to the text representation. The authors are upfront that all tasks are text-expressible, so the negative conclusion is appropriately scoped.\n\nWhere I think the paper gets shaky is the feature-survival analysis. The 88% feature-loss number comes from comparing a contextual prompt to the bare concept name, not to the kind of text passage used in the task-level evaluation. So it measures the effect of dropping context, not the effect of serializing a latent representation into natural language. The lost features might simply be context-dependent features that a full passage would also preserve. The augmentation experiment inherits this issue: the 'lost features' are the ones absent from the bare-name encoding, and showing they don't help on top of a passage doesn't prove they encode surface form. A random-feature control wouldn't fully fix this, but it would at least test whether the degradation is specific to those features or just a misalignment artifact.\n\nThis matters because the paper's strongest claim—that text destroys only surface form—is not supported for the text channel that actually matters. The task-level negative result stands on its own, but the mechanistic interpretation is weaker than the paper suggests. Also, the C1 gap (99.4 vs 80.4) has a 24.4% standard deviation on the text side and no significance test; I'd want that reported properly.\n\nOverall, this is a useful and honest negative result with a mostly sound framework. The limitation section is candid. The paper needs revision to either align the feature-survival analysis with the text-passage channel or scale back the claim about surface form. I'd send it to peer review—the framework and the cross-architecture result deserve a serious look, even if the interpretation needs tightening.\n\nBest,","headline":"A solid framework for measuring latent vs. text communication, but the headline claim about surface-form loss rests on a mismatch between the feature-survival analysis and the actual text channel.","tokens_in":13696,"tokens_out":2129,"would_cite":false,"duration_ms":19873,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Latent channels between LLM agents match text on cross-lingual concepts but never exceed it; text serialization destroys 88% of sparse features — surface form, not meaning.","keywords":["latent communication","multi-agent systems","sparse autoencoders","text serialization","representational alignment","Procrustes alignment","feature survival analysis","LLM agents"],"falsifier":"Run a communication task where the sender must convey information that cannot be faithfully serialized into text — e.g., a spatial layout, a novel visual shape, or a multi-hop reasoning state — and check whether a dense or sparse latent channel beats a text channel on a held-out receiver. A positive result would refute the paper's central negative claim. The paper itself proposes the cleanest version: align a pure vision model (no text training) to a language receiver via CCA and test whether a pathology image retrieves the correct text concept; above-chance retrieval would show latent geometr","tokens_in":12820,"feed_emoji":"🧠","tokens_out":8720,"duration_ms":76579,"temperature":0.7,"pith_summary":"The paper asks whether LLM agents can communicate more effectively by passing internal latent vectors directly instead of text, and it returns a deliberately negative answer for current tasks. It shows that text serialization destroys 88% of the Sparse Autoencoder features active in contextual encoding, replacing them with a different feature set rather than attenuating them. The decisive follow-up is that these lost features are surface form — tokenization artifacts, prompt structure, positional effects — not concept-discriminating semantics: re-injecting them degrades performance. On concept identification, attribute/sense disambiguation, and cross-lingual tasks, a latent channel matches text at best and trails by 3–10 percentage points otherwise; only bandwidth efficiency (28x compression at 99.4% probe fidelity) favors the sparse latent channel. The paper also provides clear evidence for geometric convergence between separately trained models, with a simple rotation aligning two 8B-parameter models' activation spaces to 92% top-1 concept retrieval.","feed_headline":"88% of LLM features die in text — the meaning survives","feed_subtitle":"Text already carries the signal in LLM-agent talk; sparse latent codes still compress 28x at 99% fidelity.","key_machinery":"The load-bearing object is the Sparse Autoencoder (SAE) feature dictionary — a 131,072-feature overcomplete decomposition of the 4,096-dimensional residual-stream hidden state, with about 70 active features per forward pass — used as a measurement instrument to compare two encoding paths for the same concept. The three communication channels are the dense latent channel (the full sender hidden state, 65,536 bits), the SAE-sparse channel (active feature indices and magnitudes, ~2,300 bits), and the text channel (the concept name or passage re-encoded by the receiver). Cross-architecture communication is implemented by Orthogonal Procrustes alignment — a pure rotation fitted on paired anchor c","core_discovery":"The central discovery is that text communication between LLM agents loses information at the feature level, but the lost information is not task-relevant. Using a text round-trip (a concept elicited by a contextual prompt versus the same concept encoded from its bare name), the paper finds that 88.3% of SAE features do not survive, with Jaccard overlap of only 5.8% and SAE-space cosine of 0.187. The loss is identity replacement: lost and surviving features have statistically indistinguishable activation magnitudes, so text moves the representation to a different neighborhood of feature space rather than weakening the original features. Augmentation experiments then show the lost features, wh","pith_inferences":["Editorial inference: the paper's negative conclusion is bounded by its task battery; a task requiring the sender to convey a non-textual relation (spatial layout, novel visual shape, unverbalized reasoning state) could break the ceiling, and the paper itself names cross-modal and deeper-than-language knowledge as the natural next probes.","Editorial inference: the paper's preliminary chain-of-thought results — pre-answer hidden states recover 33.1% accuracy on 'reverse-gap' reasoning tasks while text CoT reaches 96.2% — suggest latent states do carry reasoning signal; the decisive experiment would be a reasoning task where serializing the reasoning into tokens itself loses information.","Editorial inference: the finding that instruction-tuned Mistral aligns slightly worse than base Mistral hints that RLHF pushes representations away from the shared geometric core; comparing base versus chat models across more families would test whether the alignment tax is caused by instruction tuning specifically.","Editorial inference: the attention-sink injection failure in instruction-tuned models means practical latent communication cannot simply splice a vector into any layer; deployment will require template-aware injection positions, a constraint the paper's geometric-comparison evaluation sidesteps."],"forward_implications":["Current multi-agent systems that communicate in plain text are not leaving task-relevant information on the table: on all text-expressible tasks tested, text matches or beats direct latent transfer.","The SAE-sparse channel's 28x compression at 99.4% probe fidelity gives bandwidth- and interpretability-constrained deployments a practical reason to use sparse latent codes even when text remains the accuracy baseline.","Cross-architecture latent transfer is viable: a single orthogonal rotation maps one model's concept space into another's with 92% top-1 retrieval, so downstream agents could share latent states without joint training.","The 3–10pp latent-to-text gap on tasks is attributable to alignment quality, not channel fidelity or data size; better alignment (e.g., low-rank CCA, more anchors or prompts) would likely close the gap without granting latent channels an advantage on current task types.","Latent features add no value on top of a rich text passage, and the specifically 'lost' features actively hurt when re-injected — so text augmentation with hidden states is not a promising direction for concept-identification workloads."],"fun_headline_variants":["Text strips 88% of LLM-agent features, but meaning survives","LLM-agent text loses features, not task meaning","Sparse latents compress 28x, yet text matches them on tasks","Lost LLM features are surface form, not semantics","Text destroys 88% of LLM features—semantics persist"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The negative conclusion rests on the assumption that the tested tasks — concept identification, disambiguation, and cross-lingual naming, all with text-expressible answers — are representative of real multi-agent communication needs; a task whose required information cannot be serialized to tokens would void it.","fun_headline_variants_meta":{"raw":{"variants":["Text strips 88% of LLM-agent features, but meaning survives","LLM-agent text loses features, not task meaning","Sparse latents compress 28x, yet text matches them on tasks","Lost LLM features are surface form, not semantics","Text destroys 88% of LLM features—semantics persist"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000239,"raw_usage":{"total_tokens":1414,"prompt_tokens":873,"completion_tokens":541,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":617,"completion_tokens_details":{"reasoning_tokens":453}},"tokens_in":617,"tokens_out":541,"duration_ms":5285,"temperature":1.0,"reasoning_tokens":453,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-02T14:46:14.149272+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run a communication task where the sender must convey information that cannot be faithfully serialized into text — e.g., a spatial layout, a novel visual shape, or a multi-hop reasoning state — and check whether a dense or sparse latent channel beats a text channel on a held-out receiver. A positive result would refute the paper's central negative claim. The paper itself proposes the cleanest version: align a pure vision model (no text training) to a language receiver via CCA and test whether a pathology image retrieves the correct text concept; above-chance retrieval would show latent geometr","supporting_citations":[],"review_version":1}