{"id":"8d3b268b-b8fb-4653-ae04-3394de4d5954","arxiv_id":"2501.15740","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":7.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"Chalmers proposes propositional interpretability, interpreting AI in terms of beliefs, desires, and credences, and sets the challenge of thought logging all such attitudes over time.","lead":"This philosophy paper argues that AI interpretability should focus on propositional attitudes, like beliefs and goals, not just concepts, and proposes \"thought logging\" as a central challenge. It analyzes current methods like probing and sparse auto-encoders to see how close they come to this goal.","discovery_kind":"review","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Thought logging lacks a determinate target until generalized propositional attitudes are given principled psychosemantic conditions, which the paper explicitly defers.","rationale":"The most load-bearing assumption is exactly the one the Reader identified: generalized propositional attitudes must be definable precisely enough to serve as the target of thought logging, and this requires a psychosemantic theory. The paper itself acknowledges both the missing definition and the absence of a complete psychosemantics, so the concern is not manufactured. My assessment agrees with the Reader's conditional verdict: the conceptual contribution is real and the paper is honest about its open commitments, but the central challenge cannot yet be evaluated as a fully specified research program. The recommended verdict remains CONDITIONAL because the paper's own stated incompleteness is the reason for the condition, and no new evidence or argument in the stress-test pass would move the verdict in either direction.","tokens_in":19607,"tokens_out":3040,"duration_ms":34246,"concrete_test":"Take a minimal AI system with a fully known computational description and a fixed environment, such as a tabular Q-learning agent in a grid world or the Othello-GPT model from Section 7.2. Derive its propositional-attitude log twice: once using an information-based psychosemantic theory (e.g., causal/informational content) and once using a use-based theory (e.g., inferential role or a decision-theoretic representation theorem). If the two derivations disagree on at least one proposition or attitude type for the same internal state, the thought-logging target is underdetermined without further theoretical commitments. If the logs coincide across a battery of such cases, the underdetermination concern is substantially defused.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central challenge is well-posed only if there is a principled criterion for which internal states count as propositional attitudes of which contents. Section 3 defines generalized propositional attitudes partly by stipulation, saying 'I count ... even these non-sentential states as generalized propositional attitudes,' and concedes that 'it takes some work to precisely define the conditions for these or any other propositional attitudes.' Section 6 then concedes that 'we do not yet have anything close to a complete psychosemantic theory' and that 'it isn't obvious that such a theory is possible.' Since the thought-logging target is a list of propositional attitudes derived from computational and environmental facts, the absence of such a theory means the target is underdetermined: two interpreters using the information-based versus use-based psychosemantic principles described in Section 6 can disagree about whether a given activation represents p or q, or whether it is a belief-like or desire-like attitude. The resulting log would then reflect the interpreter's chosen psychosemantic framework rather than a fact about the system. This is not an objection to the importance of propositional interpretability, but it is a load-bearing gap in the paper's central research challenge.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper argues that propositional interpretability—interpreting an AI system's mechanisms and behavior in terms of propositional attitudes (beliefs, desires, credences, and their generalizations)—should be a central goal of interpretability research, complementing or extending concept-level and feature-level interpretation. It develops a taxonomy of interpretability varieties, connects the project to radical interpretation in the philosophy of mind, and proposes 'thought logging' as the central research challenge: a meta-system that takes algorithmic and environmental facts as input and outputs the AI system's propositional attitudes over time. It then assesses causal tracing, probing, sparse auto-encoders, and chain-of-thought methods as partial routes to thought logging, and replies to objections concerning AI minds, externalism, unreliability, safety, and ethics.","tokens_in":19787,"tokens_out":5814,"duration_ms":57131,"significance":"If the programmatic claim succeeds, the paper would help reorient interpretability research from locating concepts or features toward attributing propositional attitudes with specified contents and attitude types, and it would connect AI interpretability to psychosemantics, decision theory, and the philosophy of mind. The paper's strengths are its clear conceptual taxonomy, its explicit mapping of radical interpretation onto computational interpretation, its honest naming of limitations in current methods (supervision, fragility, ground truth, attitude coverage, faithfulness), and its identification of concrete subprojects such as reason logging and mechanism logging. It makes no empirical predictions and provides no machine-checked artifacts, so its contribution is conceptual and agenda-setting rather than experimental. The main risk is that the central challenge is underspecified: thought logging needs a determinate target class of attitudes and contents, and the paper explicitly defers the principled conditions for that class.","major_comments":[{"comment":"The thought-logging target proposed in §5 is a list of an AI system's propositional attitudes, but §3 introduces generalized propositional attitudes partly by stipulation ('I count ... even these non-sentential states as generalized propositional attitudes') and explicitly defers their precise conditions ('it takes some work to precisely define the conditions for these or any other propositional attitudes'). Section 6 then concedes that 'we do not yet have anything close to a complete psychosemantic theory' and that 'it isn't obvious that such a theory is possible.' The target is therefore underdetermined: information-based and use-based psychosemantic principles can disagree about whether a given activation represents p or q, or whether it is belief-like or desire-like, so the resulting log would reflect the interpreter's chosen framework rather than a fact about the system. The manuscript should either provide a working definition of the generalized attitudes to be logged, or explicitly recast thought logging as a research program with a defined target class and an acceptance criterion that does not presuppose a complete psychosemantics.","section":"§3, §6, §8"},{"comment":"The feasibility argument for thought logging is strictly conditional: 'suppose that one day we do [have a complete psychosemantic theory] ... Then we ought to be able ... to determine the system's propositional attitudes.' The paper gives no positive reason to expect the antecedent and concedes that a complete theory may be impossible. Since §5 claims that at least partial progress should be possible, the manuscript needs a separate argument or a concrete demonstration—for example, logging propositional attitudes in a small, fully observable system under an explicitly stated partial psychosemantics—to support the claim that thought logging is a tractable research program rather than a merely hypothetical one.","section":"§6"}],"minor_comments":[{"comment":"The manuscript contains unresolved draft placeholders: footnote 13 says 'A paragraph on radical vs nonradical interpretation is needed here,' and footnotes 15 and 21 promise a 'new section 8' that does not exist in this form; these must be resolved before publication.","section":"Footnotes 13, 15, 21"},{"comment":"There are several typos: 'I count will count' in §3, 'noher method' and 'on to probe' in §7.2, 'Ths raises' in §7.3, and 'albeitly' and 'journal's' in §8.","section":"§3, §7.2, §7.3, §8"},{"comment":"The example log would be clearer if the attitude names were explicitly marked as generalized propositional attitudes and if 'Judge' as a label for a credence ascription were defined on first use.","section":"§5"}],"recommendation":"major_revision","confidential_remarks":"To the editor: the manuscript is explicitly a rough draft with several unresolved placeholders. The underdetermination issue I raise is not a rejection of the importance of propositional interpretability; it is a demand that the central challenge be given a definite target. If the author prefers to keep the paper programmatic, the claims about the possibility of thought logging should be correspondingly weakened or reframed as an open research question."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The thing to know: this is not a new method paper, it is a conceptual reframing, and it is a good one. Chalmers gives interpretability a target that the field mostly gestures at—understanding what an AI system believes, desires, and assigns probability to—and gives it a name, propositional interpretability. He then sets a concrete long-term challenge, thought logging, and honestly assesses how current methods (probing, sparse auto-encoders, chain-of-thought) measure up. The Lewis framing is not decoration; it converts a vague goal into a precise format: given the computational and environmental facts, solve for the attitudes. That is genuinely useful and could organize research.\n\nThe paper does several things well. The distinction between conceptual and propositional interpretability is sharp and overdue. The survey of methods is fair, with named limitations, not strawmen. The discussion of psychosemantics is careful about what is known and what is missing. The paper also flags its own incompleteness: it is explicitly a rough draft, with a missing paragraph in Section 4 and a planned additional section. That is a real issue for peer review, but not a conceptual one.\n\nThe soft spot is exactly what the stress-test note identifies, and I think the note is right: thought logging lacks a determinate target until generalized propositional attitudes get principled psychosemantic conditions. The paper concedes this directly, both in Section 3 (defining generalized attitudes partly by stipulation) and Section 6 (no complete psychosemantic theory, and unclear that one is possible). So the central challenge is well-posed as a research program, but underdetermined as a specification. Two interpreters using different psychosemantic principles could log different attitudes from the same computational facts. That is not fatal—the paper explicitly frames this as part of the project—but it is load-bearing. I would not want a referee to miss it. At the same time, I would not demand the paper solve psychosemantics; that would be unreasonable for a foundational proposal.\n\nMinor quibbles: the Othello and mini-world discussions are familiar but used correctly. The citation pattern looks fine, engaging both the interpretability and philosophy of mind literatures. No circular derivations here; this is a conceptual argument, not a fitted model.\n\nWho is this for? AI interpretability researchers who want a philosophical grounding, and philosophers working on mental content in AI. A serious referee should engage with it, because the framing could shape the field's agenda even if the details remain open. My recommendation: send it to peer review, conditional on the author completing the missing sections. The core idea is sound and the limitations are honestly stated; the main work for revision is tightening the definition of the target, not defending the enterprise.","headline":"A clear, useful philosophical reframing of interpretability around propositional attitudes; the thought-logging challenge is real but its target stays underspecified until psychosemantics delivers.","tokens_in":779,"tokens_out":1117,"would_cite":true,"duration_ms":22128,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The core of AI interpretability should be propositional attitudes — what a system believes, desires, or gives a probability to — and the central challenge is 'thought logging.'","keywords":["propositional interpretability","thought logging","propositional attitudes","generalized propositional attitudes","mechanistic interpretability","psychosemantics","sparse auto-encoders","chain of thought"],"falsifier":"Run two independent interpretation teams through the same thought-logging exercise on the same trained model, using the paper's own criteria: psychosemantic information and use principles, behavioral validation, and the log-entry format. If the teams' attitude attributions diverge while all computational facts are agreed upon, or if the logged beliefs and goals do not predict the system's behavior better than raw activations do, then the claim that propositional attitudes are the central explanatory frame — and that thought logging is a well-defined goal — would be falsified in a concrete and observable way.","tokens_in":19396,"feed_emoji":"🧠","tokens_out":18060,"duration_ms":138048,"temperature":0.7,"pith_summary":"The paper argues that the central thing to explain about an AI system is not which concepts or features are active but what it believes, wants, and assigns probability to — its propositional attitudes. Just as human action is understood through beliefs, desires, and credences, AI systems will need to be understood the same way: knowing that a model represents a dangerous outcome is very different from knowing it has that outcome as a goal. The paper's concrete research challenge is 'thought logging': building a meta-system that records all of an AI's relevant propositional attitudes over time, ideally with reasons and mechanisms attached. It then assesses today's main interpretability methods — causal tracing, probing with classifiers, sparse auto-encoders, and chain of thought — as partial, incomplete routes toward that goal, and defends the program against objections ranging from 'AI systems have no attitudes' to externalism about content. A sympathetic reader should care because the paper reorients interpretability: the target becomes a continuous log of attitude-to-proposition attributions rather than a static map of features.","feed_headline":"Make AI readable by logging its beliefs, goals, and probabilities","feed_subtitle":"Knowing whether a system believes a danger exists or wants it to happen is the difference that matters.","key_machinery":"The load-bearing object is the propositional attitude: a relation between a system and a proposition, where the attitude can be belief, desire, credence, intention, supposition, or an engineered 'generalized' variant that does not require a mind. The paper's key move is to treat goals, world models, and probabilities — terms already applied to AI without controversy — as generalized propositional attitudes, sidestepping debates about whether AI systems have minds. Two further mechanisms carry the argument. The first is the radical-interpretation schema: given the physical or computational facts, solve for the propositional attitudes; thought logging is the engineering embodiment of that schema, with each log entry taking the form that a system has a given attitude, to a given degree, at a given time, toward a given proposition. The second is psychosemantics, the thesis that a state's content is fixed by information (what typically causes it) and use (what it drives), which supplies the in-principle bridge from algorithmic facts to attitude attributions and serves as the paper's lens for assessing each current method.","core_discovery":"The paper's central claim is that propositional interpretability — interpreting a system's mechanisms and behavior in terms of propositional attitudes, or their engineered generalizations — is a distinct and central goal of AI interpretability, and that its central challenge is thought logging: creating systems that log all of the relevant propositional attitudes of an AI system over time. The strongest assertion is that propositional attitudes are the central way we interpret and explain human beings, and are likely to be central in AI too. The argument proceeds by separating propositional from conceptual interpretability (we need to know not just that the concepts 'kill' and 'humans' are active, but whether the system is representing 'kill humans' or 'don't kill humans', and whether that representation is a belief, a desire, a credence, or a supposition), by grounding the program in radical interpretation (given the computational facts, solve for the attitudes), and by appealing to psychosemantics — theories of how a state's content is fixed by information and use — as the in-principle route from algorithmic facts to logged attitudes. On this view, thought logging is a long-term project that existing methods approach only partially: causal tracing and probing are supervised and belief-centered, sparse auto-encoders are open-ended but yield features and concepts rather than propositions and attitudes, and chain of thought outputs are pre-interpreted but unfaithful and limited to systems that actually think out loud.","pith_inferences":["A testable extension the paper leaves implicit: a propositional log's quality can be scored by whether its attitude attributions pass the belief-desire-action test — do logged beliefs and goals predict the system's next actions better than the raw activations do — and by whether probing, intervention, and behavioral methods converge on the same attributions.","The framework suggests that goal logging is the most tractable first target, because goals have a distinctive behavioral signature (persistence across counterfactual interventions), while belief attributions are where indeterminacy pressure will concentrate.","Applied to current models, the 'generalized propositional attitudes' move predicts that many so-called hallucinations are better described as unstable, prompting-relative credences or in-between believing rather than as false beliefs — a reinterpretation of existing behavioral data that the paper discusses only in outline.","If the thesis is right, interpretability benchmarking shifts from feature localization to log faithfulness and completeness, with leaderboards scoring attitude logs against ground-truth attributions in controlled environments with known goals."],"forward_implications":["If propositional interpretability is the right target, then knowing a system's concepts is not enough: interpretability efforts must attribute full attitudes (belief, desire, credence, supposition) to specific propositions, because the same represented content can be a goal or a mere model with opposite consequences.","Thought logging becomes a concrete, multi-decade benchmark: a meta-system that outputs a running list of an AI's relevant propositional attitudes, ideally extended with reason logging and mechanism logging.","Each existing method is placed by what it can and cannot log: causal tracing and probing are supervised and belief-centered; sparse auto-encoders are open-ended but log features and concepts; chain of thought yields pre-interpreted propositional text but is unfaithful and restricted to chain-of-thought systems.","Safety and ethics applications follow directly: distinguishing 'the model believes a dangerous outcome occurred' from 'the model desires that outcome' is the difference between detecting a prediction and detecting a goal.","In principle, a complete psychosemantic theory combined with near-full knowledge of a system's algorithmic state would determine its propositional attitudes, making thought logging possible in principle."],"supporting_citations":[{"why":"Supplies the canonical radical-interpretation format — attitude, degree, time, proposition — that the paper adopts for thought-log entries.","marker":"Lewis 1974"},{"why":"Sets out radical interpretation as the program of solving for beliefs, desires, and meanings; the frame the paper extends from humans to AI.","marker":"Davidson 1973"},{"why":"Names psychosemantics and frames the project of giving physical conditions for propositional attitudes, the in-principle route from algorithmic facts to logged attitudes.","marker":"Fodor 1987"},{"why":"The causal-tracing (ROME) case study that localizes and edits a belief-like fact in GPT-J, the paper's main example of use-based propositional attribution.","marker":"Meng et al 2022a"},{"why":"Trains probes to decode propositional truth-values in a mini-world, the first probing case the paper analyzes.","marker":"B. Li et al 2021"},{"why":"Decodes Othello board-state propositions from network activity, the paper's main evidence that networks carry propositional world models.","marker":"K. Li et al 2023"},{"why":"Compositional binding-based propositional probes that bind concepts into propositions; the paper's most promising route from concept logging to thought logging.","marker":"Feng et al 2024"},{"why":"Scaling Monosemanticity on Claude 3: the sparse-auto-encoder evidence for open-ended feature and concept logging, which is strong but stops short of propositions and attitudes.","marker":"Templeton et al 2024"},{"why":"The unfaithful chain-of-thought results showing models give false reasons for their answers, the key limitation of chain-of-thought as self-interpretation.","marker":"Turpin et al 2023"},{"why":"Argues we can have near-full knowledge of an AI system's algorithmic facts, the premise that makes thought logging a real-life rather than merely ideal project.","marker":"Olah 2021"}],"fun_headline_variants":["Log AI's beliefs to truly read its mind","AI interpretability needs thought logging","Propositional attitudes: the key to AI explainability","Want to understand AI? Log its propositional attitudes","Thought logging: the next step for AI transparency"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"Everything rests on the premise that a 'generalized propositional attitude' can be defined precisely enough to serve as a logging target, so that we can say — without begging the question about whether AI systems have minds — what a given system believes, wants, or assigns a probability to.","fun_headline_variants_meta":{"raw":{"variants":["Log AI's beliefs to truly read its mind","AI interpretability needs thought logging","Propositional attitudes: the key to AI explainability","Want to understand AI? Log its propositional attitudes","Thought logging: the next step for AI transparency"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000237,"raw_usage":{"total_tokens":1533,"prompt_tokens":998,"completion_tokens":535,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":614,"completion_tokens_details":{"reasoning_tokens":465}},"tokens_in":614,"tokens_out":535,"duration_ms":5599,"temperature":1.0,"reasoning_tokens":465,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-10T13:57:38.946188+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run two independent interpretation teams through the same thought-logging exercise on the same trained model, using the paper's own criteria: psychosemantic information and use principles, behavioral validation, and the log-entry format. If the teams' attitude attributions diverge while all computational facts are agreed upon, or if the logged beliefs and goals do not predict the system's behavior better than raw activations do, then the claim that propositional attitudes are the central explanatory frame — and that thought logging is a well-defined goal — would be falsified in a concrete and observable way.","supporting_citations":[],"review_version":1}