{"paper":{"title":"Dissociating Decodability and Causal Use in Bracket-Sequence Transformers","license":"http://arxiv.org/licenses/nonexclusive-distrib/1.0/","headline":"Transformers use attention to the true top-of-stack position causally for long-distance hierarchical accuracy, while decodable residual stream signals play little causal role.","cross_cats":["cs.LG"],"primary_cat":"cs.CL","authors_text":"Aryan Sharma, Cutter Dawes, Shivam Raval","submitted_at":"2026-04-24T00:26:34Z","abstract_excerpt":"When trained on tasks requiring an understanding of hierarchical structure, transformers have been found to represent this hierarchy in distinct ways: in the geometry of the residual stream, and in stack-like attention patterns maintaining a last-in, first-out ordering. However, it remains unclear whether these representations are causally used or merely decodable. We examine this gap in transformers trained on the Dyck language (a formal language of balanced bracket sequences), where the hierarchical ground truth is explicit. By probing and intervening on the residual stream and attention pat"},"claims":{"count":4,"items":[{"kind":"strongest_claim","text":"Masking attention to the true top-of-stack position causes a sharp drop in long-distance accuracy, while ablating low-dimensional residual stream subspaces has comparatively little effect.","source":"verdict.strongest_claim","status":"machine_extracted","claim_id":"C1","attestation":"unclaimed"},{"kind":"weakest_assumption","text":"That the chosen interventions (attention masking and subspace ablation) isolate the causal contribution of the probed variables without introducing unrelated side-effects on the model's computation.","source":"verdict.weakest_assumption","status":"machine_extracted","claim_id":"C2","attestation":"unclaimed"},{"kind":"one_line_summary","text":"In Dyck-language transformers, attention patterns causally use top-of-stack information while residual-stream depth and distance signals are decodable yet causally inert.","source":"verdict.one_line_summary","status":"machine_extracted","claim_id":"C3","attestation":"unclaimed"},{"kind":"headline","text":"Transformers use attention to the true top-of-stack position causally for long-distance hierarchical accuracy, while decodable residual stream signals play little causal role.","source":"verdict.pith_extraction.headline","status":"machine_extracted","claim_id":"C4","attestation":"unclaimed"}],"snapshot_sha256":"92e9a177599539cc5c9a786b4359f17d53d5d4bcaeb17b85a44ce9cece5c25e7"},"source":{"id":"2604.22128","kind":"arxiv","version":2},"verdict":{"id":"0141989a-c3b7-4b86-b69e-c9e545518087","model_set":{"reader":"grok-4.3"},"created_at":"2026-05-08T12:08:19.515606Z","strongest_claim":"Masking attention to the true top-of-stack position causes a sharp drop in long-distance accuracy, while ablating low-dimensional residual stream subspaces has comparatively little effect.","one_line_summary":"In Dyck-language transformers, attention patterns causally use top-of-stack information while residual-stream depth and distance signals are decodable yet causally inert.","pipeline_version":"pith-pipeline@v0.9.0","weakest_assumption":"That the chosen interventions (attention masking and subspace ablation) isolate the causal contribution of the probed variables without introducing unrelated side-effects on the model's computation.","pith_extraction_headline":"Transformers use attention to the true top-of-stack position causally for long-distance hierarchical accuracy, while decodable residual stream signals play little causal role."},"integrity":{"clean":false,"summary":{"advisory":1,"critical":0,"by_detector":{"doi_compliance":{"total":1,"advisory":1,"critical":0,"informational":0}},"informational":0},"endpoint":"/pith/2604.22128/integrity.json","findings":[{"note":"DOI in the printed bibliography is fragmented by whitespace or line breaks. A longer candidate (10.1162/tacla) was visible in the surrounding text but could not be confirmed against doi.org as printed.","detector":"doi_compliance","severity":"advisory","ref_index":2,"audited_at":"2026-05-20T00:14:48.227911Z","detected_doi":"10.1162/tacla","finding_type":"recoverable_identifier","verdict_class":"incontrovertible","detected_arxiv_id":null}],"available":true,"detectors_run":[{"name":"ai_meta_artifact","ran_at":"2026-05-21T11:36:17.774844Z","status":"completed","version":"1.0.0","findings_count":0},{"name":"doi_compliance","ran_at":"2026-05-20T00:14:48.227911Z","status":"completed","version":"1.0.0","findings_count":1}],"snapshot_sha256":"618146fd58359486edfab5129184d115081bf51c4534e7fcb4981b359e942a2a"},"references":{"count":0,"sample":[],"resolved_work":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57","internal_anchors":0},"formal_canon":{"evidence_count":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"author_claims":{"count":0,"strong_count":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"builder_version":"pith-number-builder-2026-05-17-v1"}