{"id":"f3baae6c-2f87-49d8-b516-1d8198475f5a","arxiv_id":"2411.16262","paper_version":1,"verdict":"REJECT","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"high","formal_verification":"none","parameter_count":3,"one_line_summary":"Probe classifiers can decode an RL agent's position from its LSTM activations in MiniHack, but this supports spatial information encoding, not a world model or self model.","lead":"This paper trains reinforcement-learning agents in a simplified video game and shows that the agents' internal network activity can be read out to reveal where the agent is located. The authors present this as evidence for rudimentary world models and a step toward machine consciousness, but the evidence is much narrower than the conclusion.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Position decoding is not a world model: the paper's central claim depends on an interpretive leap that Table III cannot support, and no control rules out sensory or temporal confounds.","rationale":"The reader's REJECT verdict is well founded. The paper's operational definition of a world model is a fitted probe's ability to decode current coordinates, which conflates 'information present' with 'model of structures and dynamics.' The authors themselves use hedging in the Discussion and explicitly list self-model/world-model distinction as future work, which undercuts the abstract's demonstrative tone. The probe results are plausible and not internally contradictory, but they are not sufficient for the central claim. As a stress-test, I looked for a way the claim could survive despite the interpretive gap: for example, if the LSTM state supported prediction of future positions rather than only current readout, that would be evidence of dynamic world-model content. The paper reports no such test. The single most decisive check is therefore a next-state prediction probe with proper episode-holdout and an untrained-agent control. If that check fails, the appropriate verdict remains REJECT; if it succeeds, the paper would still need a self-model measurement and a connection to consciousness before the central claim could be accepted. No change to the reader's verdict is needed.","tokens_in":10203,"tokens_out":5730,"duration_ms":54839,"concrete_test":"Train a probe to predict the agent's position at t+1 from the LSTM hidden/cell state at t and the action taken, with training and test episodes fully disjoint and no time-step overlap; compare against the same probe trained on an untrained agent's states. If next-position accuracy is at chance or matches the untrained-agent control while current-position accuracy is high, the representation contains a positional code but no predictive world model, invalidating the central claim. If next-position accuracy is robustly above both controls, the world-model interpretation gains real support.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim—that an RL agent 'can form rudimentary world and self models,' suggesting 'a pathway toward developing machine consciousness'—rests entirely on the step in Figure 2: if a probe trained on LSTM activations predicts the agent's current coordinates above chance, then 'the necessary information is contained in the activations. Thus the agent developed a world model.' That inference does not follow. The experiment measures only that a supervised classifier can read out a single spatial label from a recurrent state. A world model, as the authors define it, must contain 'essential structures and dynamics' that support prediction and planning; no predictive or structural content is tested. Accuracy on the random map (59.7% vs. 7.7% chance in Table III) is equally consistent with a positional code: the LSTM can store the recent trajectory and current location as task-relevant variables without modeling unobserved structure or dynamics. The self-model claim is even weaker: no internal/homeostatic variable is measured; predicting external coordinates is at most self-location, not a self model. The paper itself concedes in the Discussion that the first experiment is ambiguous and defers the self/world distinction to future work, yet the abstract asserts the conclusion as demonstrated. An additional confound is that the train/test split is not described as episode-stratified, so temporally autocorrelated LSTM states could let the probe exploit time-neighbor similarity rather than a generalizable code. Even if every reported accuracy is correct, the central claim is unsupported.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This manuscript applies linear and nonlinear probe classifiers to the activations of PPO agents trained in MiniHack rooms, with the goal of predicting the agent's current coordinates. Across three experiments—full-map input, 5x5 centered crop, and 3x3 centered crop—the probes achieve above-chance accuracies, such as 59.7% versus a 7.7% chance level in Table III. The authors interpret these results as evidence that the agent has formed rudimentary world and self models, and they link this to Damasio's theory of core consciousness and a possible pathway to machine consciousness.","tokens_in":10421,"tokens_out":3482,"duration_ms":36072,"significance":"The probe methodology is a standard and potentially useful tool, and the paper measures a real decodability signal: the agent's position is indeed linearly and nonlinearly decodable from LSTM activations. The strength is that the probes are evaluated on held-out data and the study spans multiple environment variants. However, the significance is limited by a large gap between the operational result and the interpretive claim. Above-chance coordinate decoding is expected for a task where position is reward-relevant, and it does not establish the existence of a world model in the sense the paper itself defines, namely a representation containing essential structures and dynamics for prediction and planning. The self-model claim is even weaker because no internal or homeostatic variable is probed. The paper could be a starting point for mechanistic studies of spatial encoding in RL agents, but as presented the central conclusion is not supported by the evidence.","major_comments":[{"comment":"The inference illustrated in Figure 2—that above-chance probe accuracy for the agent's current position implies 'the agent developed a world model'—is not valid. A supervised probe can read out a single spatial feature from recurrent activations without the network storing the essential structures and dynamics required by the paper's own definition of a world model. The random-map accuracy of 59.7% versus a 7.7% chance level is equally consistent with the LSTM maintaining a task-relevant positional code of its recent trajectory and current location. No predictive or structural content is tested, so the central claim goes beyond what the data support.","section":"Figure 2 and Table III"},{"comment":"The abstract states that the agent 'can form rudimentary world and self models,' but the experiments probe only external coordinates; no internal state variable such as hitpoints, resources, or reward-derived feelings is measured, so at most a form of self-location is demonstrated, not a self model. The Discussion itself concedes that 'an agent's ability to discern its position may suggest basic core consciousness, but this is not conclusive' and that future research must distinguish self-models from world models. These concessions directly contradict the definitive language of the abstract and the statement in the Discussion that the findings 'robustly confirm' a world model.","section":"Abstract and Discussion"},{"comment":"The manuscript does not describe the train/test split as episode-stratified or temporally blocked. LSTM hidden and cell states are highly autocorrelated across consecutive time steps; if samples from the same episode appear in both the training and test sets, a probe can exploit time-neighbor similarity rather than learn a generalizable position code. The paper should state whether entire episodes were held out, and ideally report probe accuracy on held-out episodes or with temporal blocking. This is a load-bearing issue for the decodability result itself, not merely for the world-model interpretation.","section":"Results, dataset construction"},{"comment":"The results are reported as single accuracies without confidence intervals, significance tests, or information about the number of training seeds. Because chance levels differ across tables due to excluded edge rows and because the paper interprets differences among maps and between hidden and cell states, the absence of variability measures makes the comparisons unreliable. At minimum, the authors should provide multiple seeds with error bars or permutation-based chance intervals, and a statistical test for the 'significantly higher than chance' claim in Figure 2.","section":"Tables I-III"}],"minor_comments":[{"comment":"The language should be hedged to match the evidence; 'demonstrate' and 'robustly confirm' are too strong for single accuracies without statistical support.","section":"Abstract and Discussion"},{"comment":"There are minor typesetting issues, such as the use of 'T' both as the episode length and in the summation, and the rendering of umlauts in author names appears corrupted in places.","section":"Methods, Equations (1)-(3)"},{"comment":"The probe training details are incomplete: the manuscript should specify input normalization, optimization details for each experiment, the number of probe parameters, and the exact procedure for splitting the 230,000 samples into 200,000 training and 30,000 test instances.","section":"Methods, Probes"},{"comment":"The paragraph referring to citation [25] is poorly integrated; it is unclear how that citation supports or contrasts with the preceding distinction between self-models and world models.","section":"Discussion"}],"recommendation":"reject","confidential_remarks":"The empirical measurements are plausible, but the conceptual leap from decodability to world and self models is the central claim, and it is not defensible from the reported experiments. The paper's own Discussion contains the caveats that undercut the abstract. A resubmission would need substantially new evidence—for example, predictive probes, intervention experiments, or tests of structural dynamics—rather than local revisions, to support the stated contribution."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Short take: the probe measurements are probably genuine, but the interpretive claim built on them is not. What is actually new is the specific application: probing LSTM activations of a MiniHack RL agent to decode its coordinates. That is a legitimate, if routine, extension of existing probing work. The experimental design is thoughtful—progressively shrinking the observation crop and adding an LSTM is a sensible way to push the agent toward reliance on memory, and the accuracies (e.g., 59.7% against a 7.7% chance baseline on the random map) are above chance and worth reporting. The paper cites the relevant probing and world-model literature, so it is not uninformed.\n\nThe soft spots are exactly where the stress-test note lands. The inference from 'a probe can classify current coordinates' to 'the agent has a world model' does not follow. By the paper's own definition, a world model must encode structures and dynamics that support prediction and planning; none of that is tested. A recurrent network can store the recent trajectory and current location as task-relevant variables without modeling unobserved structure. The self-model claim is even weaker: no internal or homeostatic variable is measured, so predicting external coordinates is at best self-location, not a self-model. The paper itself concedes in the Discussion that the first experiment is ambiguous and defers the self/world distinction to future work, yet the abstract asserts the conclusion as demonstrated. That mismatch is the core flaw.\n\nThere are also fixable but real methodological gaps: no error bars or repeated seeds, no code release, and the train/test split is not described as episode-stratified. Without that stratification, temporally autocorrelated LSTM states could let the probe exploit time-neighbor similarity rather than a generalizable code, which would weaken even the narrow decoding claim.\n\nWho this is for: readers interested in probing RL agents for spatial representations might find the raw numbers useful as a data point. Readers looking for evidence about machine consciousness will not find it here. The paper deserves a serious referee because the empirical observation is real and the topic is provocative, but it needs heavy revision—either scale the claims back to 'position is decodable' or add experiments that test predictive or structural content. My recommendation: send it to peer review with firm guidance to fix the framing and add the missing controls.","headline":"Position decoding is real, but the leap from probe accuracy to 'rudimentary world and self models' is unsupported; this is a modest interpretability result dressed as a consciousness finding.","tokens_in":11031,"tokens_out":2046,"would_cite":false,"duration_ms":20588,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper claims that a reinforcement-learning agent trained to navigate a virtual dungeon spontaneously forms rudimentary world and self models, evidenced by probes that decode the agent's coordinates from its LSTM activations with…","keywords":["machine consciousness","core consciousness","world model","self model","probing classifiers","reinforcement learning","LSTM","NetHack"],"falsifier":"Train the same probe pipeline on an agent whose LSTM is replaced by a feedforward network with the same crop input, or shuffle the order of observations so that no temporal integration is possible; if probe accuracy remains at the same level, the position signal comes from single observations, not from a memory-based world model. Alternatively, test whether a probe trained on one map can decode positions on a novel map at the same accuracy, which would indicate a general internal model rather than a map-specific code.","tokens_in":9922,"feed_emoji":"🧠","tokens_out":6597,"duration_ms":155576,"temperature":0.7,"pith_summary":"The paper tries to establish that an artificial agent trained only by reinforcement learning to play a game in a virtual environment develops preliminary forms of a world model and a self model as a byproduct of its task. The evidence comes from probes, small classifiers trained on the agent's neural activations, which can predict the agent's current grid position with accuracy well above chance. A sympathetic reading is that the agent's recurrent memory stores its location, a necessary ingredient for the kind of core consciousness described in the paper's theoretical framework. The authors are careful to say this is preliminary and not a claim that the agent is conscious.","feed_headline":"Game agent's memory encodes its own location","feed_subtitle":"Probes decode a reinforcement-learning agent's position from its LSTM, a step toward a world model.","key_machinery":"The central object is the probe, a small feedforward classifier trained on the activation vector of a single layer of a trained network to predict a property of interest. Here, each probe takes the LSTM's hidden or cell state, a 512-dimensional vector, and outputs a score for each of the 225 possible grid coordinates; the readout is accurate far above chance. The LSTM is the load-bearing component, because with only a small crop as input, reconstructing the agent's location requires memory of past observations, which is precisely the kind of internal representation a world model would need.","core_discovery":"The paper's central claim is that an RL agent playing MiniHack, a lightweight version of NetHack, ends up representing its own spatial position in the hidden and cell states of its LSTM even though it is never told its coordinates. In three experiments, probes trained on those memory states achieved accuracies of roughly 25% to 67% at predicting the agent's x and y position, against chance levels of 6.7% to 9.1%. The highest accuracies appeared when the agent's visual input was reduced to a 3x3 crop, forcing it to integrate observations over time. The authors interpret this above-chance decodability as evidence of a rudimentary world model, which they take as a step toward the core consciousness described by the theory they use.","pith_inferences":["Above-chance probe accuracy alone does not prove a world model; a positional code in the LSTM could yield the same decoding without representing environmental structure. A control that tests whether the probe generalizes to unseen maps, or whether the representation supports planning, would separate a code from a model.","The paper's discussion of successor representations suggests a connection: the same discount factor that makes the RL objective convergent also shapes expected future state occupancy, so the LSTM's hidden state may literally encode a predictive map rather than just the current location.","The same probing approach could be extended to transformers or to agents with explicit internal-state inputs such as hit points or resources, which would test the self-model half of the theory.","A sharper falsification would edit the supposed position code in the hidden state and show that the agent's behavior changes as if it believed it was elsewhere, following the activation-editing style used in the Othello-lineage work the paper cites."],"forward_implications":["If the claim holds, supposedly model-free reinforcement learning can still produce implicit internal models as a side effect of maximizing reward.","The probing pipeline offers a quantitative, task-agnostic way to look for primitive world models in artificial agents, using only their activations and ground-truth position.","The differences between hidden and cell states, and between linear and nonlinear probes, suggest the cell state carries a slightly richer positional code that may merit further study.","The paper's distinction between a stable self (the agent is always at the crop's center) and variable external observations gives a concrete direction for separating self models from world models in future experiments.","The authors note that the current environments are too simple to test for richer forms of consciousness, so follow-up work should use more complex environments and architectures."],"supporting_citations":[{"why":"Supplies the theoretical definition of core consciousness as the integration of a self model and a world model, the framework the paper tests.","marker":"[6]"},{"why":"Argues that this consciousness theory is applicable to AI systems, motivating the use of reinforcement learning agents as test subjects.","marker":"[7]"},{"why":"Provides the methodological precedent of using probes to reveal an emergent world model in a trained network, directly inspiring the probe approach.","marker":"[11]"},{"why":"Introduces linear classifier probes, the core technique used to read out positional information from the agent's activations.","marker":"[13]"},{"why":"Provides the NetHack Learning Environment, the game environment in which the agent is trained and evaluated.","marker":"[15]"},{"why":"Provides the MiniHack maps and the sandbox used for the three experimental configurations with different obstacles and observation crops.","marker":"[16]"},{"why":"Supplies the RLlib PPO implementation used to train the agents, making the specific training runs reproducible.","marker":"[24]"}],"fun_headline_variants":["LSTM probes read RL agent's position from memory","Game AI's memory maps its own coordinates","Reinforcement learning agent encodes own position in LSTM","Above-chance decoding of agent's location from neural states","Probes show game AI tracks its own position"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The argument rests on the premise that a classifier's ability to decode the agent's position from its memory is evidence of an internal world model, rather than just a convenient encoding that emerged without representing anything about the world.","fun_headline_variants_meta":{"raw":{"variants":["LSTM probes read RL agent's position from memory","Game AI's memory maps its own coordinates","Reinforcement learning agent encodes own position in LSTM","Above-chance decoding of agent's location from neural states","Probes show game AI tracks its own position"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.001099,"raw_usage":{"total_tokens":4544,"prompt_tokens":860,"completion_tokens":3684,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":476,"completion_tokens_details":{"reasoning_tokens":3609}},"tokens_in":476,"tokens_out":3684,"duration_ms":66370,"temperature":1.0,"reasoning_tokens":3609,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T13:18:56.223403+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Train the same probe pipeline on an agent whose LSTM is replaced by a feedforward network with the same crop input, or shuffle the order of observations so that no temporal integration is possible; if probe accuracy remains at the same level, the position signal comes from single observations, not from a memory-based world model. Alternatively, test whether a probe trained on one map can decode positions on a novel map at the same accuracy, which would indicate a general internal model rather than a map-specific code.","supporting_citations":[{"cited_title":"Consciousness: An overview of the phe- nomenon and of its possible neural basis,","cited_arxiv_id":null,"evidence_quote":"Supplies the theoretical definition of core consciousness as the integration of a self model and a world model, the framework the paper tests."},{"cited_title":"Will we ever have conscious machines?","cited_arxiv_id":null,"evidence_quote":"Argues that this consciousness theory is applicable to AI systems, motivating the use of reinforcement learning agents as test subjects."},{"cited_title":"Understanding intermediate layers using linear classifier probes,","cited_arxiv_id":null,"evidence_quote":"Introduces linear classifier probes, the core technique used to read out positional information from the agent's activations."},{"cited_title":"The NetHack Learning Envi- ronment,","cited_arxiv_id":null,"evidence_quote":"Provides the NetHack Learning Environment, the game environment in which the agent is trained and evaluated."},{"cited_title":"Minihack the planet: A sandbox for open-ended reinforcement learning research,","cited_arxiv_id":null,"evidence_quote":"Provides the MiniHack maps and the sandbox used for the three experimental configurations with different obstacles and observation crops."},{"cited_title":"RLlib: Abstractions for distributed reinforcement learning,","cited_arxiv_id":null,"evidence_quote":"Supplies the RLlib PPO implementation used to train the agents, making the specific training runs reproducible."}],"review_version":1}