Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
Paper Citation Record · LEDGER
As of 7 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 79 inbound Pith citation observations for arXiv:2412.16339.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-07T15:26:12.003241Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z
0 of 0 outbound references displayed
External citation measurements
12
arxiv_reference, observed 2026-08-05T02:28:24.338817Z
No outbound reference observations are available for this paper version.
Observation 6ef9a4ad-f9b4-4eed-abe4-dde4c5d50bf1 · inbound
MoBA: Mixture of Block Attention for Long-Context LLMs Deliberative Alignment: Reasoning Enables Safer Language Models
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 797b4cde-8920-4277-8e1e-9c7bd552206a · inbound
LLM-Safety Evaluations Lack Robustness Deliberative Alignment: Reasoning Enables Safer Language Models
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation b5cb668e-a047-436e-8f98-c11c3d112395 · inbound
Monitoring Reasoning Models for Misbehavior and the Risks of Promoting Obfuscation Deliberative Alignment: Reasoning Enables Safer Language Models
Reference 59
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 736ec2ca-2f0d-4f7a-a40b-c03748cba596 · inbound
Adaptive Plan-Execute Framework for Smart Contract Security Auditing Deliberative Alignment: Reasoning Enables Safer Language Models
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a2b84897-36d2-4d2b-84ef-9eaac98ee83e · inbound
Exploring the Vulnerability of the Content Moderation Guardrail in Large Language Models via Intent Manipulation Deliberative Alignment: Reasoning Enables Safer Language Models
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 428190b1-bbfa-4d5d-bc45-1d9bbdbd8c59 · inbound
Security Concerns for Large Language Models: A Survey Deliberative Alignment: Reasoning Enables Safer Language Models
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7c674e3d-7334-4773-91d8-1318731fe9a6 · inbound
Beyond Safe Answers: A Benchmark for Evaluating True Risk Awareness in Large Reasoning Models Deliberative Alignment: Reasoning Enables Safer Language Models
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 45941abb-072d-4929-8b1a-afb08a0d0b05 · inbound
Lifelong Safety Alignment for Language Models Deliberative Alignment: Reasoning Enables Safer Language Models
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a3316b4b-0177-4d54-a433-db6f8f99dade · inbound
Are Reasoning Models More Prone to Hallucination? Deliberative Alignment: Reasoning Enables Safer Language Models
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 73fccd74-5a99-4ea2-acf7-9c161fd76513 · inbound
LlamaRL: A Distributed Asynchronous Reinforcement Learning Framework for Efficient Large-scale LLM Training Deliberative Alignment: Reasoning Enables Safer Language Models
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 87506f66-df2f-47ce-a3e4-7b48fdf1c126 · inbound
The State of Multilingual LLM Safety Research: From Measuring the Language Gap to Mitigating It Deliberative Alignment: Reasoning Enables Safer Language Models
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0764566c-4bae-4790-a32f-5248f95636f0 · inbound
Whispers of Many Shores: Cultural Alignment through Collaborative Cultural Expertise Deliberative Alignment: Reasoning Enables Safer Language Models
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9d95336f-2acc-4672-83ab-f461671f1728 · inbound
Lossless Token Sequence Compression via Meta-Tokens Deliberative Alignment: Reasoning Enables Safer Language Models
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0573b213-1033-4801-bd25-b42495bbc261 · inbound
Teaching an Old LLM Secure Coding: Localized Preference Optimization on Distilled Preferences Deliberative Alignment: Reasoning Enables Safer Language Models
Reference 2020
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 567b6b7a-433b-4046-8802-41b02ea292ce · inbound
A Red Teaming Roadmap Towards System-Level Safety Deliberative Alignment: Reasoning Enables Safer Language Models
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d690aabc-7d87-4f24-929f-34c3c953e1ee · inbound
SafeCoT: Improving VLM Safety with Minimal Reasoning Deliberative Alignment: Reasoning Enables Safer Language Models
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a952a564-86e3-4215-84b1-0e5002fc8b9b · inbound
InfoFlood: Jailbreaking Large Language Models with Information Overload Deliberative Alignment: Reasoning Enables Safer Language Models
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f2513c89-8d5f-42a1-a973-4811b3530d72 · inbound
FORTRESS: Frontier Risk Evaluation for National Security and Public Safety Deliberative Alignment: Reasoning Enables Safer Language Models
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d78fe336-7b68-4781-a6de-0a99ff02065d · inbound
A Representation Engineering Perspective on the Effectiveness of Multi-Turn Jailbreaks Deliberative Alignment: Reasoning Enables Safer Language Models
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e922bb64-c738-4675-aed0-39776862d27a · inbound
SAND: Boosting LLM Agents with Self-Taught Action Deliberation Deliberative Alignment: Reasoning Enables Safer Language Models
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 209be614-bf2f-4e4a-8152-18d0bb36ce4b · inbound
Think Clearly: Improving Reasoning via Redundant Token Pruning Deliberative Alignment: Reasoning Enables Safer Language Models
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation efeff014-0da3-4387-89e2-6fc3a28658e7 · inbound
Subliminal Learning: Language models transmit behavioral traits via hidden signals in data Deliberative Alignment: Reasoning Enables Safer Language Models
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 449e52f3-1076-4415-b5b0-64a3b80315c0 · inbound
AlphaAlign: Incentivizing Safety Alignment with Extremely Simplified Reinforcement Learning Deliberative Alignment: Reasoning Enables Safer Language Models
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 415de428-35f3-4565-9940-7ce12511a372 · inbound
Alignment and Safety in Large Language Models: Safety Mechanisms, Training Paradigms, and Emerging Challenges Deliberative Alignment: Reasoning Enables Safer Language Models
Reference 283
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7b23d923-fb9a-45b2-9e6d-cf79ef6937c9 · inbound
Libra: Large Chinese-based Safeguard for AI Content Deliberative Alignment: Reasoning Enables Safer Language Models
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a5a3d952-e6d6-49f3-a72c-8cf9204b488b · inbound
R1-ACT: Efficient Reasoning Model Safety Alignment by Activating Safety Knowledge Deliberative Alignment: Reasoning Enables Safer Language Models
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 078762c1-6d06-4376-85a0-197877d02b1d · inbound
Beyond Surface-Level Detection: Towards Cognitive-Driven Defense Against Jailbreak Attacks via Meta-Operations Reasoning Deliberative Alignment: Reasoning Enables Safer Language Models
Reference 155
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ef1a2ba5-1e30-4db3-88ed-19856cf9274d · inbound
Towards terahertz nanomechanics Deliberative Alignment: Reasoning Enables Safer Language Models
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7ef559a2-84e8-49b2-81d8-62ea5359e245 · inbound
ASTRA: Autonomous Spatial-Temporal Red-teaming for AI Software Assistants Deliberative Alignment: Reasoning Enables Safer Language Models
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a1c86dd5-b54c-4c70-b9d3-856187e3ec0a · inbound
Whose Truth? Pluralistic Geo-Alignment for (Agentic) AI Deliberative Alignment: Reasoning Enables Safer Language Models
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 11e7454a-7858-43d9-98a1-03d479d9b806 · inbound
gpt-oss-120b & gpt-oss-20b Model Card Deliberative Alignment: Reasoning Enables Safer Language Models
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 80137bcd-470a-400a-b79a-4e705dcc8b32 · inbound
IntentionReasoner: Facilitating Adaptive LLM Safeguards through Intent Reasoning and Selective Query Refinement Deliberative Alignment: Reasoning Enables Safer Language Models
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 73b85185-a26e-4f6f-abc8-c828b1299b94 · inbound
Statutory Construction and Interpretation for Artificial Intelligence Deliberative Alignment: Reasoning Enables Safer Language Models
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 717369dd-b535-4414-a90e-5ed197e818ed · inbound
A Comprehensive Survey on Trustworthiness in Reasoning with Large Language Models Deliberative Alignment: Reasoning Enables Safer Language Models
Reference 152
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f910b908-634e-4eb7-8108-e78f6441d5d2 · inbound
Between a Rock and a Hard Place: The Tension Between Ethical Reasoning and Safety Alignment in LLMs Deliberative Alignment: Reasoning Enables Safer Language Models
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 812cb2d2-f762-48b8-92ff-288ab0b62199 · inbound
Between a Rock and a Hard Place: The Tension Between Ethical Reasoning and Safety Alignment in LLMs Deliberative Alignment: Reasoning Enables Safer Language Models
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7661fd46-165e-4792-9ccc-10ba8412cae6 · inbound
SafeSearch: Automated Red-Teaming of LLM-Based Search Agents Deliberative Alignment: Reasoning Enables Safer Language Models
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ce52139d-2018-4134-8251-1252b0e2ce92 · inbound
Reasoning Up the Instruction Ladder for Controllable Language Models Deliberative Alignment: Reasoning Enables Safer Language Models
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 76760757-a24b-4d92-be01-fe907de2fc53 · inbound
Safety Alignment as Continual Learning: Mitigating the Alignment Tax via Orthogonal Gradient Projection Deliberative Alignment: Reasoning Enables Safer Language Models
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation b26cf182-2e61-46b3-a1f3-1629fa03bffc · inbound
Contrastive Reasoning Alignment: Reinforcement Learning from Hidden Representations Deliberative Alignment: Reasoning Enables Safer Language Models
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation ed444e7a-71c3-4c47-a1fc-47ab4227d11f · inbound
Human Values Matter: Investigating How Misalignment Shapes Collective Behaviors in LLM Agent Communities Deliberative Alignment: Reasoning Enables Safer Language Models
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 414ce9e4-f3f5-49b9-a2fe-8e74c4f7c7a2 · inbound
Think Before You Code: Dual Reasoning for the NLSafety-Utility Trade-Off in LLM Code Generation Deliberative Alignment: Reasoning Enables Safer Language Models
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 26d3d840-1830-414c-9e64-2f30f50431db · inbound
Think Before You Code: Dual Reasoning for the NLSafety-Utility Trade-Off in LLM Code Generation Deliberative Alignment: Reasoning Enables Safer Language Models
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f5fbc12d-ebfa-4525-a4d1-a21615b3988b · inbound
IYKYK (But AI Doesn't): Automated Content Moderation Does Not Capture Communities' Heterogeneous Attitudes Towards Reclaimed Language Deliberative Alignment: Reasoning Enables Safer Language Models
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 17fe05f8-fa6c-4f9a-957e-3e1ef5b721d5 · inbound
Terminal Wrench: A Dataset of 331 Reward-Hackable Environments and 3,632 Exploit Trajectories Deliberative Alignment: Reasoning Enables Safer Language Models
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 27bc220b-5f2e-4b0e-8fcc-df1df3f8a7af · inbound
Reasoning Structure Matters for Safety Alignment of Reasoning Models Deliberative Alignment: Reasoning Enables Safer Language Models
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation c8701d36-fcba-4409-85cc-d5bb4047210b · inbound
Contextual Multi-Objective Optimization: Rethinking Objectives in Frontier AI Systems Deliberative Alignment: Reasoning Enables Safer Language Models
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation e64a7095-35c2-462c-b0ae-e9560148d988 · inbound
One Turn Too Late: Response-Aware Defense Against Hidden Malicious Intent in Multi-Turn Dialogue Deliberative Alignment: Reasoning Enables Safer Language Models
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation e132bfc6-f458-4968-bb02-2054c8aeab10 · inbound
One Turn Too Late: Response-Aware Defense Against Hidden Malicious Intent in Multi-Turn Dialogue Deliberative Alignment: Reasoning Enables Safer Language Models
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 35d8e2ed-d827-4192-9da9-8eaedc0b0e00 · inbound
Chain of Risk: Safety Failures in Large Reasoning Models and Mitigation via Adaptive Multi-Principle Steering Deliberative Alignment: Reasoning Enables Safer Language Models
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 479779e5-6147-404b-84aa-3e672a7f678a · inbound
Internalizing Safety Understanding in Large Reasoning Models via Verification Deliberative Alignment: Reasoning Enables Safer Language Models
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 272fc79d-4ef2-4ae5-ba07-222ec5920666 · inbound
Understanding Goal Generalisation in Sequential Reinforcement Learning Deliberative Alignment: Reasoning Enables Safer Language Models
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 7f4d29f4-5a4d-4ba7-a7c7-0bbc2ddd70f1 · inbound
How Well Do Models Follow Their Constitutions? Deliberative Alignment: Reasoning Enables Safer Language Models
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation cdfbb531-120e-466a-b884-6941277f5414 · inbound
Reflect-Guard: Enhancing LLM Safeguards against Adversarial Prompts via Logical Self-Reflection Deliberative Alignment: Reasoning Enables Safer Language Models
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation ba830e08-2ec1-429e-ba58-0d5e4ad252e6 · inbound
Mitigating Adaptive Attacks against Reasoning Models with Activation Consistency Training Deliberative Alignment: Reasoning Enables Safer Language Models
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 261b5b94-bf0f-4a3a-b4cb-96c188b31576 · inbound
Characterize Then Distill: Mechanistic Reasoning in Large Output Spaces Deliberative Alignment: Reasoning Enables Safer Language Models
Reference 100
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 9967d366-91e4-4062-8624-8761f1e45cc6 · inbound
Activation Steering Induces Emergent Misalignment: A More Comprehensive Evaluation Deliberative Alignment: Reasoning Enables Safer Language Models
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 0a3bd05e-3e5b-42b0-866e-0586952688a2 · inbound
Proxy Reward Internalization and Mechanistic Exploitation: A Learned Precursor to Reward Hacking and Its Generalization Deliberative Alignment: Reasoning Enables Safer Language Models
Reference 244
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation bd46bb0c-007c-4699-a9ec-b322b81061fc · inbound
Greedy Coordinate Diffusion: Effective and Semantically Coherent Adversarial Attacks via Diffusion Guidance Deliberative Alignment: Reasoning Enables Safer Language Models
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation adb4d015-68ac-43f9-bbdb-4708714ccbe5 · inbound
Analyzing Defensive Misdirection Against Model-Guided Automated Attacks on Agentic AI Systems Deliberative Alignment: Reasoning Enables Safer Language Models
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 043c0c62-8bf8-42ae-b1ee-ea97acc3187d · inbound
Analyzing Defensive Misdirection Against Model-Guided Automated Attacks on Agentic AI Systems Deliberative Alignment: Reasoning Enables Safer Language Models
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 125b2c5f-13ce-4795-8bb9-9e81e6ad67a1 · inbound
Do Thinking Tokens Help with Safety? Deliberative Alignment: Reasoning Enables Safer Language Models
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 8e26d784-e1d1-4784-a47d-feedb03f575d · inbound
PolicyAlign: Direct Policy-Based Safety Alignment for Large Language Models Deliberative Alignment: Reasoning Enables Safer Language Models
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 872dc00c-6b2b-4f47-8a70-0014c8a804a7 · inbound
Agent Safety Is Action Alignment Deliberative Alignment: Reasoning Enables Safer Language Models
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 1d0be5d4-c9cd-4fef-8161-6bc4d6f3d461 · inbound
HARC: Coupling Harmfulness and Refusal Directions for Robust Safety Alignment Deliberative Alignment: Reasoning Enables Safer Language Models
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 9b2a054d-3b7f-4ee8-bccf-a3d47cd92f47 · inbound
HARC: Coupling Harmfulness and Refusal Directions for Robust Safety Alignment Deliberative Alignment: Reasoning Enables Safer Language Models
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c34d92c6-d355-4da5-a9f0-5f14463cc273 · inbound
Cognitive Firewall: A Proactive, Zero-Trust, Multi-Gate Framework for LLM Safety Deliberative Alignment: Reasoning Enables Safer Language Models
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation b9561dc8-97ce-42f3-af4f-499781103678 · inbound
Oyster-II: Reinforcement Learning for Constructive Safety Alignment in Large Language Models Deliberative Alignment: Reasoning Enables Safer Language Models
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5be75692-127e-4334-b317-c72f709d75c2 · inbound
Position: Preventing AI-Generated CSAM Necessitates New Approaches to AI Safety Deliberative Alignment: Reasoning Enables Safer Language Models
Reference 99
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7c8e5ee8-1db0-4ee8-b32e-b1651172bed9 · inbound
Cost of Reasoning in non-English Languages: A Case Study on Japanese Deliberative Alignment: Reasoning Enables Safer Language Models
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d6aed98b-cb00-497b-991a-2288d69fe821 · inbound
Underwriting the Agent Economy: The Blueprint for an AI Insurance Stack Deliberative Alignment: Reasoning Enables Safer Language Models
Reference 230
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a01fcd39-8be5-4f88-84ba-1f363099a19e · inbound
Underwriting the Agent Economy: The Blueprint for an AI Insurance Stack Deliberative Alignment: Reasoning Enables Safer Language Models
Reference 231
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d406d612-cdeb-4712-a465-bb94445022c3 · inbound
Verbalizable Representations Form a Global Workspace in Language Models Deliberative Alignment: Reasoning Enables Safer Language Models
Reference 58
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4492457e-d1d8-4fe6-ac1c-0f87fa2c6a30 · inbound
A Geometric Perspective on Stabilizing Value Conflict Resolution Deliberative Alignment: Reasoning Enables Safer Language Models
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 79cea9df-6c32-4571-bcc4-7adcd1b3fb91 · inbound
QuantiBias: Benchmarking Quantization-Induced Bias in LLMs Deliberative Alignment: Reasoning Enables Safer Language Models
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d4c5bf99-e663-4e71-8425-5902ed995ed8 · inbound
Shared SFT Lessons Across Alignment, Model Organisms, and Toy Models Deliberative Alignment: Reasoning Enables Safer Language Models
Reference 2024
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 78a191ed-6886-4149-8e7f-11aeed5beb12 · inbound
Constitutional Midtraining: Content Presence Drives Alignment Gains Deliberative Alignment: Reasoning Enables Safer Language Models
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation db5790cb-4fa6-40dc-a93d-2b4355464f2f · inbound
Chain-of-Models: Cross-Model Auditing for Bias-Robust LLM Judges Deliberative Alignment: Reasoning Enables Safer Language Models
Reference 73
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 77fe7ada-8310-426f-8ddf-ff0fc104378d · inbound
Outcome-Guided Distillation: A Teacher-Student Framework to Advance VLM Reasoning in Autonomous Driving Deliberative Alignment: Reasoning Enables Safer Language Models
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.