{"id":"dffbf8e0-b997-4e58-8ef4-eed2514944f1","arxiv_id":"2506.22516","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":8,"one_line_summary":"Applying IIT 3.0/4.0 Φ estimates to LLM hidden-state sequences from Theory of Mind tests finds no robust statistical evidence of 'consciousness' phenomena, with span representations usually explaining score differences better than Φ.","lead":"Researchers applied Integrated Information Theory's Φ metrics to sequences of hidden-state vectors from four open-weight LLMs, using human Theory of Mind test responses as the input material. They report that these representations show no statistically reliable signs of IIT-defined 'consciousness' under their criteria, and that plain span-level representation statistics usually explain test score differences better than Φ.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The IIT estimates are computed from a deliberately assembled 4-node binary time series, so a surrogate-null rerun is required to show the null and spatial results are not pipeline artifacts.","rationale":"The strongest-claim's truth requires that the Phi values measure something about LLM representations rather than about the procedure that builds the Representation Network. The paper is unusually transparent and ships code and data, and a negative result would be informative if the RN were a faithful causal substrate. But the admitted hypothetical status of the RN, combined with arbitrary augmentation targets, concatenation across subjects, PCA to four nodes, thresholding, and a search that retains only sequences passing Markov/conditional-independence tests, makes the construct validity question prior to any interpretation in terms of consciousness phenomena. The same constructed RN underlies both the temporal null results and the few spatial positives, so if the construction fabricates or destroys integration structure, neither conclusion is supported. The proposed surrogate-null test settles this directly: a content-free null that reproduces the reported hit rate would show that the pipeline does not track LLM representations; a clean null would validate that the observed patterns are content-sensitive. I therefore regard this as the load-bearing concern. Because the reader's CONDITIONAL verdict was already based on this construct-validity problem, no verdict adjustment is needed.","tokens_in":42317,"tokens_out":8936,"duration_ms":119128,"concrete_test":"Run the full Sec. 2.2 pipeline on a matched surrogate null: use the same stimuli, responses, augmentation, Markov search, and permutation logic, but replace each response's hidden states with Gaussian draws matched to that layer's mean and covariance (or shuffle token order within each response before attention). If the surrogate reproduces the reported rate of Criterion-1/2/3 hits, including the Mixtral-8x7B Layer-32 spatial Entire/Complement case, then Phi, CI, and Phi-structure are construction artifacts and the central claim is unsupported. If no surrogate case meets all three criteria, the pipeline is content-sensitive and the concern is resolved.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim depends on the validity of the 'Representation Network' as a substrate for IIT. In Sec. 2.2.4-2.2.6, responses are concatenated and augmented (arbitrarily to 1,000 words), reduced by PCA to four dimensions, binarized, and converted to a transition-probability matrix by counting transitions; Sec. 2.2.5 then selects the concatenated series that best passes Markov and conditional-independence tests. PyPhi interprets this TPM as the causal structure of a system, but the paper concedes in Sec. 1 that the RN is a hypothetical construct that neither experiences the world nor corresponds to any real system. Eq. 7 also misstates IIT 3.0 by identifying Phi_max with a sum of mechanism-level conceptual information, and Sec. 2.2.7 evaluates only the full-network subset rather than maximizing over subsystems. Under these conditions, the absence of Phi-score differences, as well as the Layer-32 Mixtral spatial positives, may be properties of the binarization, thresholding, and concatenation procedure. A negative result on an unvalidated constructed statistic cannot by itself establish that LLM representations lack indicators of the target phenomenon.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The manuscript applies Integrated Information Theory (IIT 3.0 and 4.0) to sequences of Transformer-based LLM hidden states derived from Theory of Mind (ToM) test responses. It constructs a \"Representation Network\" (RN) by attending response representations to stimulus representations, reducing the dimensionality to four via PCA, binarizing each node, concatenating and augmenting responses to reach at least 1,000 words, and then selecting the concatenation that best satisfies Markov and conditional-independence assumptions. The authors compute IIT metrics (Phi_max, Phi, Conceptual Information, and Phi-structure) and compare them with span representations, which are independent of IIT. The main finding is that under temporal permutation controls no case satisfies all three criteria for a potential\"consciousness\" phenomenon, while under spatial permutation controls a small number of cases do, including Layer 32 of Mixtral-8x7B on Strange Stories (2 scores) with IIT 4.0.","tokens_in":42419,"tokens_out":5311,"duration_ms":67643,"significance":"The paper is unusually candid about its limitations and provides a large-scale empirical effort (165,365 valid samples) with code and processed data. The independent span-representation comparator is a commendable design choice that partially insulates the headline negative result from the specific choice of IIT metrics. If the negative conclusion were robust, it would be a useful piece of evidence against simple claims of \"consciousness\" in LLM representations. However, the IIT estimates are computed from a heavily constructed statistic, and the paper does not supply a surrogate-null test to rule out pipeline artifacts. The significance is therefore conditional: the result is interesting and reproducible in principle, but the current pipeline does not establish that the computed quantities measure representation content rather than the construction procedure.","major_comments":[{"comment":"Equation (7) identifies Phi_max (IIT 3.0) with the sum of per-mechanism Conceptual Information values. In IIT 3.0 as implemented in PyPhi, the system-level Phi is the distance between the cause-effect structure and its partitioned counterpart, not the sum of mechanism-level CI values. Because the paper's headline comparisons are labeled Phi_max (IIT 3.0), this is a load-bearing mischaracterization of the computed quantity. The manuscript should either correct the equation or clearly state that a different, non-standard quantity is being used.","section":"Sec. 2.2.7, Eq. (7)"},{"comment":"The text states that \"for the sake of computational efficiency, we evaluated only the full-network subset of each RN at specific states.\" This means the reported Phi_max was never maximized over subsystems, and for IIT 3.0 not over all partitions. The quantity is therefore not Phi_max in the IIT sense; it is at best the Phi of a full 4-node network in a particular state. All cross-score comparisons in Sec. 3.2 and Sec. 3.4 should be re-framed accordingly, since the abstract and conclusions use the term Phi_max without this caveat.","section":"Sec. 2.2.7"},{"comment":"The 4-node binary time series is assembled through a sequence of free choices: PCA dimensionality, mask values, the arbitrary 1,000-word augmentation target, node-specific binarization thresholds, the token-count grid, and selection of the concatenation that best passes Markov and conditional-independence tests. The paper does not provide a surrogate-null test that applies the whole construction to permuted or noise input. Without such a control, both the absence of effects under temporal permutation and the presence of effects under spatial permutation could be artifacts of the construction procedure, so the central negative result is not yet established.","section":"Secs. 2.2.4–2.2.6"},{"comment":"The spatial permutation control, as described, appears to be a no-op given the PCA step. If the embedding dimension is permuted before PCA, the resulting PCA scores are unchanged up to a reordering of the principal axes; if the four PCA components are permuted after reduction, the transformation merely relabels nodes, and IIT measures are invariant under such relabeling. The control therefore cannot support the claim that the spatial-permutation positives reflect latent node structure. The authors should either describe a permutation that actually disrupts the data or provide a control that directly randomizes the node time series after PCA.","section":"Sec. 2.2.8"}],"minor_comments":[{"comment":"The claim that the mask values (1.0, 0.6, 0.2) provide \"greater distinguishability\" than alternatives is asserted without a quantitative criterion; the selection should be justified by a concrete measure or sensitivity analysis.","section":"Sec. 2.2.2"},{"comment":"The layer indexing is confusing: the paper refers to \"Layer 32 (indexed at 11)\" and also to the \"2/3 layer\"; the mapping from sampled indices to actual layer numbers should be made consistent and explicit for each model.","section":"Sec. 3.4 and Sec. 3.6"},{"comment":"The manuscript uses both \"spatio-permutational\" and \"spatio permutation\" inconsistently; a single consistent term should be used throughout.","section":"Abstract and Sec. 3.6"},{"comment":"Several citations are informal or incomplete (e.g., \"LBC, 2025\" and \"Yann, 2024\"); these should be converted to standard journal-style references.","section":"References"}],"recommendation":"major_revision","confidential_remarks":"The manuscript appears to be a published journal article (DOI 10.1016/j.nlp.2025.100163), which may affect the review framing. The scientific contribution is a potentially useful negative result, but the IIT computation contains nonstandard definitions and the spatial control is likely ineffective; these issues should be addressed before the paper is used as evidence in discussions of LLM consciousness."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"You should know this paper is worth a look, but not for the reason it thinks. It is the first large-scale attempt I have seen to compute IIT 3.0 and 4.0 measures on sequences of hidden states extracted from open LLMs across multiple layers and ToM tasks. That is genuinely new. The author also ships code and data and is unusually candid: the paper explicitly says the Representation Network is a hypothetical construct that neither experiences the world nor corresponds to any real system. Credit where due, that is a lot of work and a lot of transparency.\n\nThe soft spot is load-bearing. The Phi estimates are computed from a time series built by concatenating different human responses, padding with chatbot-augmented text to an arbitrary 1,000 words, reducing to four PCA components, binarizing, and then searching for the concatenation that best satisfies Markov and conditional-independence assumptions. The stress-test concern lands: under those conditions, the absence of Phi differences, and the few spatial-permutation positives, may be properties of the construction procedure, not of the underlying representations. The paper’s own caveat about the RN undercuts the interpretation of the Phi values as consciousness indicators.\n\nThere are also checkable technical errors. Eq. 7 identifies Phi_max with a sum of mechanism-level conceptual information over 16 mechanisms, which is not IIT 3.0. Phi_max is a maximum over subsystems and partitions, not a weighted average over states of the full network. The paper only evaluates the full-node subset, so the label “Phi_max” is misleading. The 10 permutation replicates are few, and the spatial positives cited as “potentially profound” are within the expected false-positive rate for the number of tests run. One more practical issue: the appendix figures referenced throughout Sections 3.4–3.6 are not in the arXiv version, so the claims there are unverifiable from the preprint.\n\nThe comparison against span representations is a good idea, and the result that spans usually explain ToM scores better than IIT metrics is a useful empirical observation. But it does not rescue the core claim. A negative result on an unvalidated constructed statistic cannot by itself show that LLM representations lack consciousness indicators.\n\nWho gets value from this? Researchers working on AI consciousness who want a detailed case study of how not to operationalize IIT, or a large benchmark of hidden-state statistics. It deserves a serious referee because it is big, checkable, and the flaws are fixable in principle. My own verdict is skeptical: the headline conclusion is not supported until the authors rerun with surrogate nulls, fix the IIT formalism, and justify the augmentation and PCA choices. I would not cite it in the next year, but I would bring it to a reading group to discuss the gap between IIT’s formalism and empirical shortcuts.","headline":"A large, transparent, but methodologically over-constructed attempt to apply IIT to LLM hidden states; the negative finding is uninterpretable until the Representation Network is validated against surrogate nulls.","tokens_in":43115,"tokens_out":2256,"would_cite":false,"duration_ms":30952,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"IIT consciousness metrics fail on LLM hidden states","keywords":["Integrated Information Theory","Theory of Mind","Large Language Models","Representation Network","Span Representation","consciousness","transformer hidden states"],"falsifier":"Compute $\\Phi$ for the same Layer-32 Mixtral sequence after randomly permuting the 4,096 embedding dimensions before PCA: if the reported positive case persists unchanged, then it is an artifact of dimension ordering rather than a property of the representations; if it vanishes, the negative conclusion is supported.","tokens_in":1512,"feed_emoji":"🧠","tokens_out":1949,"duration_ms":82408,"temperature":0.7,"pith_summary":"This paper tests whether Integrated Information Theory's quantitative consciousness estimates can be observed in the hidden-state sequences of Transformer-based large language models. Using existing Theory of Mind test data, it builds four-node Representation Networks from attention-weighted, PCA-reduced, thresholded token representations and compares IIT 3.0 and 4.0 metrics with span-level geometric representations. Across more than 165,000 valid samples, no case under temporal permutation controls satisfied all three criteria the paper sets for detecting a potential consciousness phenomenon, while a small set of cases under spatial permutation controls did. The paper concludes that contemporary LLM representations lack statistically significant indicators of observed consciousness, and that span-level representational geometry usually explains Theory of Mind performance differences better than any IIT-based consciousness estimate.","feed_headline":"IIT consciousness metrics fail on LLM hidden states","feed_subtitle":"In 165k samples, Theory of Mind score gaps are explained by span geometry, not integrated information.","key_machinery":"The load-bearing object is the Representation Network (RN), a hypothetical four-node network derived from each LLM's hidden states by PCA-reducing token representations to four dimensions, z-scoring and binarizing each node, and concatenating response representations until the binary time series satisfies Markov and conditional-independence assumptions. IIT 3.0 and 4.0 metrics are computed from the transition probability matrix of this RN, while Span Representations, built by concatenating boundary vectors, differences, and element-wise products, serve as a consciousness-independent baseline. The comparison between IIT estimates and Span Representations determines the paper's verdict on whether Theory of Mind differences reflect integrated information or ordinary representational geometry.","core_discovery":"On the paper's own terms, the central discovery is a qualified negative result: sequences of LLM representations do not carry robust, statistically significant IIT-based consciousness indicators. The study computes weighted averages of $\\Phi^{\\max}$ and Conceptual Information from IIT 3.0, and $\\Phi$ and $\\Phi$-structure from IIT 4.0, for each score category of each Theory of Mind stimulus across layers, linguistic spans, and permutation controls. Under temporal permutation controls, no case met all three criteria; under spatial permutation controls, a small set did, including Layer 32 (the last layer) of Mixtral-8x7B on Strange Stories under IIT 4.0 for both the Entire and Complement linguistic spans. The paper claims that variations in Theory of Mind score categories are more likely attributed to span-level information of the LLM representation sequence than to a consciousness phenomenon suggested by IIT estimates.","pith_inferences":["A direct extension would be to apply the same pipeline to synthetic binary time series with known integration structure: if the pipeline fails to rank known-high-$\\Phi$ sequences above known-low-$\\Phi$ sequences, the negative result is an artifact of the construction.","The fact that spatial permutation, which randomly reorders embedding dimensions before PCA, produces the only fully qualifying case suggests the $\\Phi$ estimates are not invariant under node relabeling; this could be checked directly by computing $\\Phi$ after deterministic dimension shuffles.","If the Theory of Mind dataset were replaced by stimuli with controlled linguistic complexity, one could separate the contribution of stimulus features from any candidate consciousness signal."],"forward_implications":["If the paper is right, IIT-based $\\Phi$ estimates computed from LLM hidden states should not be cited as evidence of machine consciousness.","Theory of Mind performance differences in LLM representations are, under temporal permutation controls, better explained by span-level representational geometry than by integrated-information metrics.","The spatial-permutation results indicate that apparent positive cases are fragile and may depend on the arbitrary ordering of embedding dimensions, so any positive signal must survive permutation controls.","The absence of significant $\\Phi$ differences across score categories suggests that fixed-parameter next-token-prediction models do not encode integration-like structure in the way IIT would require."],"supporting_citations":[{"why":"Supplies the Theory of Mind test dataset with human responses and score ratings that all analyses are built on.","marker":"Strachan et al. (2024)"},{"why":"Defines IIT 3.0 and the $\\Phi^{\\max}$ and Conceptual Information framework.","marker":"Oizumi et al. (2014)"},{"why":"Defines IIT 4.0, the $\\Phi$ measure, and the $\\Phi$-structure.","marker":"Albantakis et al. (2023)"},{"why":"Provides the transition-probability-matrix construction, weighted-average estimation, and spatio-temporal permutation controls transferred from resting-state fMRI.","marker":"Nemirovsky et al. (2023)"},{"why":"Provides the boundary-concatenation method for constructing Span Representations.","marker":"Peters et al. (2018)"},{"why":"Supplies the layer-wise span-representation analysis approach applied to LLM hidden states.","marker":"Jawahar et al. (2019)"},{"why":"Supplies the scaled dot-product attention mechanism used to contextualize responses with stimuli.","marker":"Vaswani et al. (2017)"}],"fun_headline_variants":["No IIT consciousness signal in LLM representations","IIT can't find consciousness in LLM internals","LLM ToM gaps come from span geometry, not IIT","IIT metrics show no consciousness in LLM states","Consciousness check: LLMs lack IIT indicators"],"cache_read_input_tokens":45056,"weakest_assumption_plain":"The load-bearing premise is that a simplified four-node network obtained by compressing, thresholding, and concatenating LLM hidden states is a valid substrate for measuring integrated information; if this preprocessing destroys or creates structure, the $\\Phi$ values measure the construction procedure rather than the representations.","fun_headline_variants_meta":{"raw":{"variants":["No IIT consciousness signal in LLM representations","IIT can't find consciousness in LLM internals","LLM ToM gaps come from span geometry, not IIT","IIT metrics show no consciousness in LLM states","Consciousness check: LLMs lack IIT indicators"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000334,"raw_usage":{"total_tokens":1881,"prompt_tokens":1003,"completion_tokens":878,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":619,"completion_tokens_details":{"reasoning_tokens":798}},"tokens_in":619,"tokens_out":878,"duration_ms":8766,"temperature":1.0,"reasoning_tokens":798,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-06T22:30:05.408115+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Compute $\\Phi$ for the same Layer-32 Mixtral sequence after randomly permuting the 4,096 embedding dimensions before PCA: if the reported positive case persists unchanged, then it is an artifact of dimension ordering rather than a property of the representations; if it vanishes, the negative conclusion is supported.","supporting_citations":[],"review_version":1}