{"record_type":"pith_number_record","schema_url":"https://pith.science/schemas/pith-number/v1.json","pith_number":"pith:2026:MXWQVYC4O52MSNQTI5RWB2TLRR","short_pith_number":"pith:MXWQVYC4","schema_version":"1.0","canonical_sha256":"65ed0ae05c7774c93613476360ea6b8c582bb5999ee2b9843679446a72605544","source":{"kind":"arxiv","id":"2604.08169","version":2},"attestation_state":"computed","paper":{"title":"Activation Steering for Aligned Open-ended Generation without Sacrificing Coherence","license":"http://arxiv.org/licenses/nonexclusive-distrib/1.0/","headline":"Activation steering recovers honesty and compassion in LLMs during generation without harming coherence.","cross_cats":[],"primary_cat":"cs.AI","authors_text":"Alberto Tosato, Gauthier Gidel, Martin Zborowski, Niklas Herbster, Tommaso Tosato","submitted_at":"2026-04-09T12:28:22Z","abstract_excerpt":"Alignment in LLMs is more brittle than commonly assumed: misalignment can be induced by adversarial prompts, benign fine-tuning, emergent misalignment, and goal misgeneralization. Recent evidence suggests that some misalignment behaviors are encoded as linear structure in activation space, making it tractable via activation steering, which could be used as a lightweight runtime defense. We implement three methods: Steer-With-Fixed-Coefficient (SwFC), which applies uniform additive steering, and two novel projection-aware methods, Steer-to-Target-Projection (StTP) and Steer-to-Mirror-Projection"},"verification_status":{"content_addressed":true,"pith_receipt":true,"author_attested":false,"weak_author_claims":0,"strong_author_claims":0,"externally_anchored":false,"storage_verified":false,"citation_signatures":0,"replication_records":0,"graph_snapshot":true,"references_resolved":false,"formal_links_present":false},"canonical_record":{"source":{"id":"2604.08169","kind":"arxiv","version":2},"metadata":{"license":"http://arxiv.org/licenses/nonexclusive-distrib/1.0/","primary_cat":"cs.AI","submitted_at":"2026-04-09T12:28:22Z","cross_cats_sorted":[],"title_canon_sha256":"b7ac2178cdb7bf7901ff45250307135d1c9aa1c6e0263c373383eb23cb847deb","abstract_canon_sha256":"d6bd77c30842d2b074ca98fbc6b53310614751d51a126ed3fea9a3ce870e7c61"},"schema_version":"1.0"},"receipt":{"kind":"pith_receipt","key_id":"pith-v1-2026-05","algorithm":"ed25519","signed_at":"2026-07-03T01:16:53.055528Z","signature_b64":"1T8sbz0OxOBHPIAi3PUCC3JC5iTKqPWbnuSKp1lHqqf9oOCbhaXAgut57jTOHoAYn7KH/gHUDugq+VlQ8qMBDg==","signed_message":"canonical_sha256_bytes","builder_version":"pith-number-builder-2026-05-17-v1","receipt_version":"0.3","canonical_sha256":"65ed0ae05c7774c93613476360ea6b8c582bb5999ee2b9843679446a72605544","last_reissued_at":"2026-07-03T01:16:53.054940Z","signature_status":"signed_v1","first_computed_at":"2026-07-03T01:16:53.054940Z","public_key_fingerprint":"8d4b5ee74e4693bcd1df2446408b0d54"},"graph_snapshot":{"paper":{"title":"Activation Steering for Aligned Open-ended Generation without Sacrificing Coherence","license":"http://arxiv.org/licenses/nonexclusive-distrib/1.0/","headline":"Activation steering recovers honesty and compassion in LLMs during generation without harming coherence.","cross_cats":[],"primary_cat":"cs.AI","authors_text":"Alberto Tosato, Gauthier Gidel, Martin Zborowski, Niklas Herbster, Tommaso Tosato","submitted_at":"2026-04-09T12:28:22Z","abstract_excerpt":"Alignment in LLMs is more brittle than commonly assumed: misalignment can be induced by adversarial prompts, benign fine-tuning, emergent misalignment, and goal misgeneralization. Recent evidence suggests that some misalignment behaviors are encoded as linear structure in activation space, making it tractable via activation steering, which could be used as a lightweight runtime defense. We implement three methods: Steer-With-Fixed-Coefficient (SwFC), which applies uniform additive steering, and two novel projection-aware methods, Steer-to-Target-Projection (StTP) and Steer-to-Mirror-Projection"},"claims":{"count":4,"items":[{"kind":"strongest_claim","text":"All methods substantially recover target traits (honesty and compassion) while preserving coherence. StTP and StMP better maintain general capabilities (MMLU, MT-Bench, AlpacaEval) and produce less repetition in multi-turn conversations.","source":"verdict.strongest_claim","status":"machine_extracted","claim_id":"C1","attestation":"unclaimed"},{"kind":"weakest_assumption","text":"That malicious system prompts function as a valid controlled proxy for broader misalignment triggers including adversarial prompts, benign fine-tuning, emergent misalignment, and goal misgeneralization.","source":"verdict.weakest_assumption","status":"machine_extracted","claim_id":"C2","attestation":"unclaimed"},{"kind":"one_line_summary","text":"Projection-aware activation steering using logistic regression recovers honesty and compassion under malicious prompts while preserving coherence and benchmark performance better than uniform steering.","source":"verdict.one_line_summary","status":"machine_extracted","claim_id":"C3","attestation":"unclaimed"},{"kind":"headline","text":"Activation steering recovers honesty and compassion in LLMs during generation without harming coherence.","source":"verdict.pith_extraction.headline","status":"machine_extracted","claim_id":"C4","attestation":"unclaimed"}],"snapshot_sha256":"7d4d06ef4a1c8c759134f71c8535cf84f4044a4a4e6df0a90f44af037df7f493"},"source":{"id":"2604.08169","kind":"arxiv","version":2},"verdict":{"id":"e89ed02d-3a0b-4020-89a8-a5f96373fb88","model_set":{"reader":"grok-4.3"},"created_at":"2026-05-10T17:13:39.349353Z","strongest_claim":"All methods substantially recover target traits (honesty and compassion) while preserving coherence. StTP and StMP better maintain general capabilities (MMLU, MT-Bench, AlpacaEval) and produce less repetition in multi-turn conversations.","one_line_summary":"Projection-aware activation steering using logistic regression recovers honesty and compassion under malicious prompts while preserving coherence and benchmark performance better than uniform steering.","pipeline_version":"pith-pipeline@v0.9.0","weakest_assumption":"That malicious system prompts function as a valid controlled proxy for broader misalignment triggers including adversarial prompts, benign fine-tuning, emergent misalignment, and goal misgeneralization.","pith_extraction_headline":"Activation steering recovers honesty and compassion in LLMs during generation without harming coherence."},"integrity":{"clean":true,"summary":{"advisory":0,"critical":0,"by_detector":{},"informational":0},"endpoint":"/pith/2604.08169/integrity.json","findings":[],"available":true,"detectors_run":[],"snapshot_sha256":"c28c3603d3b5d939e8dc4c7e95fa8dfce3d595e45f758748cecf8e644a296938"},"references":{"count":0,"sample":[],"resolved_work":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57","internal_anchors":0},"formal_canon":{"evidence_count":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"author_claims":{"count":0,"strong_count":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"builder_version":"pith-number-builder-2026-05-17-v1"},"aliases":[{"alias_kind":"arxiv","alias_value":"2604.08169","created_at":"2026-07-03T01:16:53.055007+00:00"},{"alias_kind":"arxiv_version","alias_value":"2604.08169v2","created_at":"2026-07-03T01:16:53.055007+00:00"},{"alias_kind":"doi","alias_value":"10.48550/arxiv.2604.08169","created_at":"2026-07-03T01:16:53.055007+00:00"},{"alias_kind":"pith_short_12","alias_value":"MXWQVYC4O52M","created_at":"2026-07-03T01:16:53.055007+00:00"},{"alias_kind":"pith_short_16","alias_value":"MXWQVYC4O52MSNQT","created_at":"2026-07-03T01:16:53.055007+00:00"},{"alias_kind":"pith_short_8","alias_value":"MXWQVYC4","created_at":"2026-07-03T01:16:53.055007+00:00"}],"events":[],"event_summary":{},"paper_claims":[],"inbound_citations":{"count":0,"internal_anchor_count":0,"sample":[]},"formal_canon":{"evidence_count":0,"sample":[],"anchors":[]},"links":{"html":"https://pith.science/pith/MXWQVYC4O52MSNQTI5RWB2TLRR","json":"https://pith.science/pith/MXWQVYC4O52MSNQTI5RWB2TLRR.json","graph_json":"https://pith.science/api/pith-number/MXWQVYC4O52MSNQTI5RWB2TLRR/graph.json","events_json":"https://pith.science/api/pith-number/MXWQVYC4O52MSNQTI5RWB2TLRR/events.json","paper":"https://pith.science/paper/MXWQVYC4"},"agent_actions":{"view_html":"https://pith.science/pith/MXWQVYC4O52MSNQTI5RWB2TLRR","download_json":"https://pith.science/pith/MXWQVYC4O52MSNQTI5RWB2TLRR.json","view_paper":"https://pith.science/paper/MXWQVYC4","resolve_alias":"https://pith.science/api/pith-number/resolve?arxiv=2604.08169&json=true","fetch_graph":"https://pith.science/api/pith-number/MXWQVYC4O52MSNQTI5RWB2TLRR/graph.json","fetch_events":"https://pith.science/api/pith-number/MXWQVYC4O52MSNQTI5RWB2TLRR/events.json","actions":{"anchor_timestamp":"https://pith.science/pith/MXWQVYC4O52MSNQTI5RWB2TLRR/action/timestamp_anchor","attest_storage":"https://pith.science/pith/MXWQVYC4O52MSNQTI5RWB2TLRR/action/storage_attestation","attest_author":"https://pith.science/pith/MXWQVYC4O52MSNQTI5RWB2TLRR/action/author_attestation","sign_citation":"https://pith.science/pith/MXWQVYC4O52MSNQTI5RWB2TLRR/action/citation_signature","submit_replication":"https://pith.science/pith/MXWQVYC4O52MSNQTI5RWB2TLRR/action/replication_record"}},"created_at":"2026-07-03T01:16:53.055007+00:00","updated_at":"2026-07-03T01:16:53.055007+00:00"}