Pith. sign in

Paper Citation Record · LEDGER

Deliberative Alignment: Reasoning Enables Safer Language Models

As of 7 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 79 inbound Pith citation observations for arXiv:2412.16339.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.16339 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 79 of 79 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 79 of 79 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T15:26:12.003241Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

12
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 6ef9a4ad-f9b4-4eed-abe4-dde4c5d50bf1 · inbound

MoBA: Mixture of Block Attention for Long-Context LLMs cites this paper.

MoBA: Mixture of Block Attention for Long-Context LLMs Deliberative Alignment: Reasoning Enables Safer Language Models

Reference 51

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T06:15:46.165967Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-16T06:15:46.085555Z digest=sha256:f9c9ad301dbcac6726cf41736402c7554c23f8a829bf0aab4359737503d0610f

Observation 797b4cde-8920-4277-8e1e-9c7bd552206a · inbound

LLM-Safety Evaluations Lack Robustness cites this paper.

LLM-Safety Evaluations Lack Robustness Deliberative Alignment: Reasoning Enables Safer Language Models

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-23T01:27:21.396795Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-23T01:26:45.402983Z digest=sha256:b2c1c169e2e92333a3a249890ebb552b261a5e2576531dfbcf366fb485defef6

Observation b5cb668e-a047-436e-8f98-c11c3d112395 · inbound

Monitoring Reasoning Models for Misbehavior and the Risks of Promoting Obfuscation cites this paper.

Monitoring Reasoning Models for Misbehavior and the Risks of Promoting Obfuscation Deliberative Alignment: Reasoning Enables Safer Language Models

Reference 59

Resolution
verified exact
arxiv_id, observed 2026-05-21T07:24:12.967449Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-21T07:24:12.845841Z digest=sha256:7ff7ceaad6259ef4d1a50f28b7286e3b0a736103321c746fd38193baed0e6478

Observation 736ec2ca-2f0d-4f7a-a40b-c03748cba596 · inbound

Adaptive Plan-Execute Framework for Smart Contract Security Auditing cites this paper.

Adaptive Plan-Execute Framework for Smart Contract Security Auditing Deliberative Alignment: Reasoning Enables Safer Language Models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T15:26:12.003241Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:26:12.003241Z digest=sha256:02698eceba8602f536a3922d56222966e1bbd45038612ff47dfe161d8f740076

Observation a2b84897-36d2-4d2b-84ef-9eaac98ee83e · inbound

Exploring the Vulnerability of the Content Moderation Guardrail in Large Language Models via Intent Manipulation cites this paper.

Exploring the Vulnerability of the Content Moderation Guardrail in Large Language Models via Intent Manipulation Deliberative Alignment: Reasoning Enables Safer Language Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T14:32:29.074711Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:32:29.074711Z digest=sha256:a096d85fe2f5bc3e69c7e47245707f4df32137ea51d593a2da67246a13964b2d

Observation 428190b1-bbfa-4d5d-bc45-1d9bbdbd8c59 · inbound

Security Concerns for Large Language Models: A Survey cites this paper.

Security Concerns for Large Language Models: A Survey Deliberative Alignment: Reasoning Enables Safer Language Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T14:27:03.474448Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:27:03.474448Z digest=sha256:90db0d78c9629d09e45ef67c8ba25dbc7002353372526644af4fddb90a874fe0

Observation 7c674e3d-7334-4773-91d8-1318731fe9a6 · inbound

Beyond Safe Answers: A Benchmark for Evaluating True Risk Awareness in Large Reasoning Models cites this paper.

Beyond Safe Answers: A Benchmark for Evaluating True Risk Awareness in Large Reasoning Models Deliberative Alignment: Reasoning Enables Safer Language Models

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T14:13:54.133916Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:13:54.133916Z digest=sha256:c6de973763dc200e54b348ddbf893588270911072efbd59890cdd771218e75d9

Observation 45941abb-072d-4929-8b1a-afb08a0d0b05 · inbound

Lifelong Safety Alignment for Language Models cites this paper.

Lifelong Safety Alignment for Language Models Deliberative Alignment: Reasoning Enables Safer Language Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T14:00:05.091365Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:00:05.091365Z digest=sha256:4b7b6b04aabf4ff6a2f51e7110beb056ff738326375af88ca4ea57541bcf085f

Observation a3316b4b-0177-4d54-a433-db6f8f99dade · inbound

Are Reasoning Models More Prone to Hallucination? cites this paper.

Are Reasoning Models More Prone to Hallucination? Deliberative Alignment: Reasoning Enables Safer Language Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:30.366747Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:45:30.366747Z digest=sha256:4b5a89d3b962ee5e509cc3577e2cf5df64d74d807206b28dcce40bfd51b9d42f

Observation 73fccd74-5a99-4ea2-acf7-9c161fd76513 · inbound

LlamaRL: A Distributed Asynchronous Reinforcement Learning Framework for Efficient Large-scale LLM Training cites this paper.

LlamaRL: A Distributed Asynchronous Reinforcement Learning Framework for Efficient Large-scale LLM Training Deliberative Alignment: Reasoning Enables Safer Language Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:25.860418Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:45:25.860418Z digest=sha256:ae61901bf909cbb532e79cce122ba5a7ee5bfe55afbece1bcb04a4365b3722ec

Observation 87506f66-df2f-47ce-a3e4-7b48fdf1c126 · inbound

The State of Multilingual LLM Safety Research: From Measuring the Language Gap to Mitigating It cites this paper.

The State of Multilingual LLM Safety Research: From Measuring the Language Gap to Mitigating It Deliberative Alignment: Reasoning Enables Safer Language Models

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T12:40:18.665783Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:40:18.665783Z digest=sha256:9b0286c84db38dabdf6cb7c17da07532e3b0a94cacab0137649d27e2f5d1bbb4

Observation 0764566c-4bae-4790-a32f-5248f95636f0 · inbound

Whispers of Many Shores: Cultural Alignment through Collaborative Cultural Expertise cites this paper.

Whispers of Many Shores: Cultural Alignment through Collaborative Cultural Expertise Deliberative Alignment: Reasoning Enables Safer Language Models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T12:13:21.052638Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:13:21.052638Z digest=sha256:c4924e658c9c6ac05b14d3d238258a3e8c4ba3b59a398ec725afa62ceb8e71ee

Observation 9d95336f-2acc-4672-83ab-f461671f1728 · inbound

Lossless Token Sequence Compression via Meta-Tokens cites this paper.

Lossless Token Sequence Compression via Meta-Tokens Deliberative Alignment: Reasoning Enables Safer Language Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T12:14:26.266557Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:14:26.266557Z digest=sha256:899c265b0564df2e0d16577f4e2d81ae3da467434fc233b0f27fcc4149a93a6b

Observation 0573b213-1033-4801-bd25-b42495bbc261 · inbound

Teaching an Old LLM Secure Coding: Localized Preference Optimization on Distilled Preferences cites this paper.

Teaching an Old LLM Secure Coding: Localized Preference Optimization on Distilled Preferences Deliberative Alignment: Reasoning Enables Safer Language Models

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-07T12:10:30.276451Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:10:30.276451Z digest=sha256:db4679526ec0f481312ea62e3e1be56b9a8154901059c7eabc8c37cb1b187425

Observation 567b6b7a-433b-4046-8802-41b02ea292ce · inbound

A Red Teaming Roadmap Towards System-Level Safety cites this paper.

A Red Teaming Roadmap Towards System-Level Safety Deliberative Alignment: Reasoning Enables Safer Language Models

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T12:11:19.235153Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:11:19.235153Z digest=sha256:e2add3c88f3e33fd1942ff4127be00984ecc36e5f666ec6eb00ae78cf088ee25

Observation d690aabc-7d87-4f24-929f-34c3c953e1ee · inbound

SafeCoT: Improving VLM Safety with Minimal Reasoning cites this paper.

SafeCoT: Improving VLM Safety with Minimal Reasoning Deliberative Alignment: Reasoning Enables Safer Language Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T05:19:00.051825Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:19:00.051825Z digest=sha256:344f8d84bb9e6881f43d2f1fa27b37d424387b99af6df0134b0ad158529bd8ed

Observation a952a564-86e3-4215-84b1-0e5002fc8b9b · inbound

InfoFlood: Jailbreaking Large Language Models with Information Overload cites this paper.

InfoFlood: Jailbreaking Large Language Models with Information Overload Deliberative Alignment: Reasoning Enables Safer Language Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T01:02:28.553458Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T01:02:28.553458Z digest=sha256:e008cbc240234e7e801781075f461d5b1f18dc3363017d2aa4159af376d783bd

Observation f2513c89-8d5f-42a1-a973-4811b3530d72 · inbound

FORTRESS: Frontier Risk Evaluation for National Security and Public Safety cites this paper.

FORTRESS: Frontier Risk Evaluation for National Security and Public Safety Deliberative Alignment: Reasoning Enables Safer Language Models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T00:15:00.611082Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:15:00.611082Z digest=sha256:24720616b1082fdc6e83e094ba7540b313db49bb53b479e50c807d80ec683c65

Observation d78fe336-7b68-4781-a6de-0a99ff02065d · inbound

A Representation Engineering Perspective on the Effectiveness of Multi-Turn Jailbreaks cites this paper.

A Representation Engineering Perspective on the Effectiveness of Multi-Turn Jailbreaks Deliberative Alignment: Reasoning Enables Safer Language Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T21:50:24.108617Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:50:24.108617Z digest=sha256:adc237ec8b103ac145c02f47838fd42cc8f61569753c99cc4a679741e4afa1ec

Observation e922bb64-c738-4675-aed0-39776862d27a · inbound

SAND: Boosting LLM Agents with Self-Taught Action Deliberation cites this paper.

SAND: Boosting LLM Agents with Self-Taught Action Deliberation Deliberative Alignment: Reasoning Enables Safer Language Models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T18:47:19.985383Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:47:19.985383Z digest=sha256:f7c0a664ea097b166ed49f48ed36e2fb2e9c77a8f6b5cd02f62c5ced5c5c7d1a

Observation 209be614-bf2f-4e4a-8152-18d0bb36ce4b · inbound

Think Clearly: Improving Reasoning via Redundant Token Pruning cites this paper.

Think Clearly: Improving Reasoning via Redundant Token Pruning Deliberative Alignment: Reasoning Enables Safer Language Models

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T00:23:17.986195Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:23:17.986195Z digest=sha256:8d8d6811d21ce39a11600499f6b7bd53c53fe41bd91a8c0c23f6d5400463a7ac

Observation efeff014-0da3-4387-89e2-6fc3a28658e7 · inbound

Subliminal Learning: Language models transmit behavioral traits via hidden signals in data cites this paper.

Subliminal Learning: Language models transmit behavioral traits via hidden signals in data Deliberative Alignment: Reasoning Enables Safer Language Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T15:53:17.551777Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:53:17.551777Z digest=sha256:65da1622f678e89ba42970259f96d5d389337ad95ffb51c1cdc8f8b6eba8f910

Observation 449e52f3-1076-4415-b5b0-64a3b80315c0 · inbound

AlphaAlign: Incentivizing Safety Alignment with Extremely Simplified Reinforcement Learning cites this paper.

AlphaAlign: Incentivizing Safety Alignment with Extremely Simplified Reinforcement Learning Deliberative Alignment: Reasoning Enables Safer Language Models

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T15:48:01.386002Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:48:01.386002Z digest=sha256:1028c43dbbdfd95cfa1dd06ef2ad7d30ef5b79950e6419fd54accfc491d17c68

Observation 415de428-35f3-4565-9940-7ce12511a372 · inbound

Alignment and Safety in Large Language Models: Safety Mechanisms, Training Paradigms, and Emerging Challenges cites this paper.

Alignment and Safety in Large Language Models: Safety Mechanisms, Training Paradigms, and Emerging Challenges Deliberative Alignment: Reasoning Enables Safer Language Models

Reference 283

Resolution
unresolved
no resolver link, observed 2026-08-06T14:13:07.297537Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:13:07.297537Z digest=sha256:26dc42d036ab9a06b022bd1998283c3677264cbea2c0ac138ebd337d43a728c6

Observation 7b23d923-fb9a-45b2-9e6d-cf79ef6937c9 · inbound

Libra: Large Chinese-based Safeguard for AI Content cites this paper.

Libra: Large Chinese-based Safeguard for AI Content Deliberative Alignment: Reasoning Enables Safer Language Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T12:17:38.663630Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:17:38.663630Z digest=sha256:01f0cd847b736446fe8fce1ebd15dec23c27a15a286b047a5c775b2c725d8b7a

Observation a5a3d952-e6d6-49f3-a72c-8cf9204b488b · inbound

R1-ACT: Efficient Reasoning Model Safety Alignment by Activating Safety Knowledge cites this paper.

R1-ACT: Efficient Reasoning Model Safety Alignment by Activating Safety Knowledge Deliberative Alignment: Reasoning Enables Safer Language Models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T10:18:24.466989Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T10:18:24.466989Z digest=sha256:d258435cea5524acdf98d9dbace9ca068fc4efd5e518ee2a5c16623e155c0eba

Observation 078762c1-6d06-4376-85a0-197877d02b1d · inbound

Beyond Surface-Level Detection: Towards Cognitive-Driven Defense Against Jailbreak Attacks via Meta-Operations Reasoning cites this paper.

Beyond Surface-Level Detection: Towards Cognitive-Driven Defense Against Jailbreak Attacks via Meta-Operations Reasoning Deliberative Alignment: Reasoning Enables Safer Language Models

Reference 155

Resolution
unresolved
no resolver link, observed 2026-08-06T04:47:25.538461Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T04:47:25.538461Z digest=sha256:ca606b211da6c4db0057b6fdab41cc00a1355db769720f7601d85d78d2d26e6b

Observation ef1a2ba5-1e30-4db3-88ed-19856cf9274d · inbound

Towards terahertz nanomechanics cites this paper.

Towards terahertz nanomechanics Deliberative Alignment: Reasoning Enables Safer Language Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T01:03:50.244026Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T01:03:50.244026Z digest=sha256:e609a4278eabe0faa1a322ce8c5381c15fe62369deca4b06f089010aad52ec7f

Observation 7ef559a2-84e8-49b2-81d8-62ea5359e245 · inbound

ASTRA: Autonomous Spatial-Temporal Red-teaming for AI Software Assistants cites this paper.

ASTRA: Autonomous Spatial-Temporal Red-teaming for AI Software Assistants Deliberative Alignment: Reasoning Enables Safer Language Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T01:04:33.591580Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T01:04:33.591580Z digest=sha256:081b97ebfa6d7560f8ad9f2a7c8c1560137b098e026cdd1730b701a8e0a9914a

Observation a1c86dd5-b54c-4c70-b9d3-856187e3ec0a · inbound

Whose Truth? Pluralistic Geo-Alignment for (Agentic) AI cites this paper.

Whose Truth? Pluralistic Geo-Alignment for (Agentic) AI Deliberative Alignment: Reasoning Enables Safer Language Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-05T23:23:31.962576Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:23:31.962576Z digest=sha256:7b2b3b545a8413cc6544d6d23c89c37eb8b7e20ae684c4b51e8edb5e93a820a4

Observation 11e7454a-7858-43d9-98a1-03d479d9b806 · inbound

gpt-oss-120b & gpt-oss-20b Model Card cites this paper.

gpt-oss-120b & gpt-oss-20b Model Card Deliberative Alignment: Reasoning Enables Safer Language Models

Reference 29

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T12:22:54.679028Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T12:22:54.633089Z digest=sha256:b52e865989be4a0414e00ee2f735efe4fe5ffef537998b8f6469c0aa2bcbfcbb

Observation 80137bcd-470a-400a-b79a-4e705dcc8b32 · inbound

IntentionReasoner: Facilitating Adaptive LLM Safeguards through Intent Reasoning and Selective Query Refinement cites this paper.

IntentionReasoner: Facilitating Adaptive LLM Safeguards through Intent Reasoning and Selective Query Refinement Deliberative Alignment: Reasoning Enables Safer Language Models

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-05T15:20:14.188064Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:20:14.188064Z digest=sha256:33f3102f4f472d66371c5088289f6f45f10eded4702d47c1f14a86ad714b7464

Observation 73b85185-a26e-4f6f-abc8-c828b1299b94 · inbound

Statutory Construction and Interpretation for Artificial Intelligence cites this paper.

Statutory Construction and Interpretation for Artificial Intelligence Deliberative Alignment: Reasoning Enables Safer Language Models

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-05T12:52:54.701047Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:52:54.701047Z digest=sha256:7ffa8e5abacafaac35ddc5e5d32c3410d52b5ea7e757a83fe639175c1461098f

Observation 717369dd-b535-4414-a90e-5ed197e818ed · inbound

A Comprehensive Survey on Trustworthiness in Reasoning with Large Language Models cites this paper.

A Comprehensive Survey on Trustworthiness in Reasoning with Large Language Models Deliberative Alignment: Reasoning Enables Safer Language Models

Reference 152

Resolution
unresolved
no resolver link, observed 2026-08-05T10:39:07.027988Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:39:07.027988Z digest=sha256:4495c56912a13385b20985084293a5594d87a508509499d851cd4da8d248d627

Observation f910b908-634e-4eb7-8108-e78f6441d5d2 · inbound

Between a Rock and a Hard Place: The Tension Between Ethical Reasoning and Safety Alignment in LLMs cites this paper.

Between a Rock and a Hard Place: The Tension Between Ethical Reasoning and Safety Alignment in LLMs Deliberative Alignment: Reasoning Enables Safer Language Models

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-18T19:36:47.435169Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-18T19:36:23.882344Z digest=sha256:03630241642c2d1d4a54b54928c9bf630a2da97c344ca3d91653bd638fbb739c

Observation 812cb2d2-f762-48b8-92ff-288ab0b62199 · inbound

Between a Rock and a Hard Place: The Tension Between Ethical Reasoning and Safety Alignment in LLMs cites this paper.

Between a Rock and a Hard Place: The Tension Between Ethical Reasoning and Safety Alignment in LLMs Deliberative Alignment: Reasoning Enables Safer Language Models

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-05T10:36:15.617560Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:36:15.617560Z digest=sha256:4424bd2892e1fe1db207fdcd5aa99613cdd8eb6e52ef4a148d58d8f7b3719cca

Observation 7661fd46-165e-4792-9ccc-10ba8412cae6 · inbound

SafeSearch: Automated Red-Teaming of LLM-Based Search Agents cites this paper.

SafeSearch: Automated Red-Teaming of LLM-Based Search Agents Deliberative Alignment: Reasoning Enables Safer Language Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-04T14:43:49.835349Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T14:43:49.835349Z digest=sha256:39c5a51de02632b004da06c2a88a81cc267ced2203f648c3e03b543ab733b97f

Observation ce52139d-2018-4134-8251-1252b0e2ce92 · inbound

Reasoning Up the Instruction Ladder for Controllable Language Models cites this paper.

Reasoning Up the Instruction Ladder for Controllable Language Models Deliberative Alignment: Reasoning Enables Safer Language Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-04T07:07:52.834565Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T07:07:52.834565Z digest=sha256:d25b7e39f8c250625e70f437144a3a9eb074bf94ffbfe13a5d58895758ecb1b4

Observation 76760757-a24b-4d92-be01-fe907de2fc53 · inbound

Safety Alignment as Continual Learning: Mitigating the Alignment Tax via Orthogonal Gradient Projection cites this paper.

Safety Alignment as Continual Learning: Mitigating the Alignment Tax via Orthogonal Gradient Projection Deliberative Alignment: Reasoning Enables Safer Language Models

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-16T06:40:42.343205Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-16T06:39:17.785715Z digest=sha256:50a72ed0596826903301d12b22b772c2260483f935d38b17a614412cb0353eef

Observation b26cf182-2e61-46b3-a1f3-1629fa03bffc · inbound

Contrastive Reasoning Alignment: Reinforcement Learning from Hidden Representations cites this paper.

Contrastive Reasoning Alignment: Reinforcement Learning from Hidden Representations Deliberative Alignment: Reasoning Enables Safer Language Models

Reference 10

Resolution
metadata mismatch
arxiv_id, observed 2026-05-21T11:24:08.445919Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-21T11:23:26.852673Z digest=sha256:fc55071b5a58d26cb4d12950d2b18b17ae36400b146b8d016f13691c8f748447

Observation ed444e7a-71c3-4c47-a1fc-47ab4227d11f · inbound

Human Values Matter: Investigating How Misalignment Shapes Collective Behaviors in LLM Agent Communities cites this paper.

Human Values Matter: Investigating How Misalignment Shapes Collective Behaviors in LLM Agent Communities Deliberative Alignment: Reasoning Enables Safer Language Models

Reference 4

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T22:25:51.877085Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T19:52:16.543718Z digest=sha256:b8b937cfc33b7f1f40758c7194e31195220c7a32f693b57917d97882fe2004cc

Observation 414ce9e4-f3f5-49b9-a2fe-8e74c4f7c7a2 · inbound

Think Before You Code: Dual Reasoning for the NLSafety-Utility Trade-Off in LLM Code Generation cites this paper.

Think Before You Code: Dual Reasoning for the NLSafety-Utility Trade-Off in LLM Code Generation Deliberative Alignment: Reasoning Enables Safer Language Models

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-11T09:56:03.213083Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T15:44:50.762366Z digest=sha256:731c8cb27ba4be4e83cc20a33c868f1e23aaaefb58a145b0940aeab287de7dfe

Observation 26d3d840-1830-414c-9e64-2f30f50431db · inbound

Think Before You Code: Dual Reasoning for the NLSafety-Utility Trade-Off in LLM Code Generation cites this paper.

Think Before You Code: Dual Reasoning for the NLSafety-Utility Trade-Off in LLM Code Generation Deliberative Alignment: Reasoning Enables Safer Language Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-03T00:19:56.880861Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T00:19:56.880861Z digest=sha256:3c9e594711d491985dc7fa0b89b5fb28afe40088d3f393194a470cb84e283dad

Observation f5fbc12d-ebfa-4525-a4d1-a21615b3988b · inbound

IYKYK (But AI Doesn't): Automated Content Moderation Does Not Capture Communities' Heterogeneous Attitudes Towards Reclaimed Language cites this paper.

IYKYK (But AI Doesn't): Automated Content Moderation Does Not Capture Communities' Heterogeneous Attitudes Towards Reclaimed Language Deliberative Alignment: Reasoning Enables Safer Language Models

Reference 39

Resolution
verified exact
arxiv_id, observed 2026-05-10T08:22:37.099023Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T08:19:50.909364Z digest=sha256:98ba4ada47fb491fa8ede045011b8d061bef2ce0943e424f698b68449febc715

Observation 17fe05f8-fa6c-4f9a-957e-3e1ef5b721d5 · inbound

Terminal Wrench: A Dataset of 331 Reward-Hackable Environments and 3,632 Exploit Trajectories cites this paper.

Terminal Wrench: A Dataset of 331 Reward-Hackable Environments and 3,632 Exploit Trajectories Deliberative Alignment: Reasoning Enables Safer Language Models

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-10T05:51:10.098655Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T05:48:44.687520Z digest=sha256:5eee59b8576ceb9e0e97230b9f77adb1ae1ba028ab7d4bda254d4e31f5c382a6

Observation 27bc220b-5f2e-4b0e-8fcc-df1df3f8a7af · inbound

Reasoning Structure Matters for Safety Alignment of Reasoning Models cites this paper.

Reasoning Structure Matters for Safety Alignment of Reasoning Models Deliberative Alignment: Reasoning Enables Safer Language Models

Reference 25

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T12:56:04.339490Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-10T02:36:57.093584Z digest=sha256:333c3ad12c3b5c526f4cdbd9a635ae6a833b4eda8301c289a6ae9b18c84b88d3

Observation c8701d36-fcba-4409-85cc-d5bb4047210b · inbound

Contextual Multi-Objective Optimization: Rethinking Objectives in Frontier AI Systems cites this paper.

Contextual Multi-Objective Optimization: Rethinking Objectives in Frontier AI Systems Deliberative Alignment: Reasoning Enables Safer Language Models

Reference 19

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T23:41:18.981416Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-07T16:26:04.152045Z digest=sha256:68382513793651caf236442e354c260754c065346ec83fce275c32670eb69cd7

Observation e64a7095-35c2-462c-b0ae-e9560148d988 · inbound

One Turn Too Late: Response-Aware Defense Against Hidden Malicious Intent in Multi-Turn Dialogue cites this paper.

One Turn Too Late: Response-Aware Defense Against Hidden Malicious Intent in Multi-Turn Dialogue Deliberative Alignment: Reasoning Enables Safer Language Models

Reference 8

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T19:41:09.007290Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-08T11:17:19.079380Z digest=sha256:44ad30a1059e3552bbf905c12f47204ac3d92fd745eaa1d4fbc967245eb17b15

Observation e132bfc6-f458-4968-bb02-2054c8aeab10 · inbound

One Turn Too Late: Response-Aware Defense Against Hidden Malicious Intent in Multi-Turn Dialogue cites this paper.

One Turn Too Late: Response-Aware Defense Against Hidden Malicious Intent in Multi-Turn Dialogue Deliberative Alignment: Reasoning Enables Safer Language Models

Reference 8

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T07:57:31.670860Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T07:53:06.928500Z digest=sha256:aa39fb2fb6ac62d0a09239b9eaee5361e946716a47b1b9b18999abe964a44eed

Observation 35d8e2ed-d827-4192-9da9-8eaedc0b0e00 · inbound

Chain of Risk: Safety Failures in Large Reasoning Models and Mitigation via Adaptive Multi-Principle Steering cites this paper.

Chain of Risk: Safety Failures in Large Reasoning Models and Mitigation via Adaptive Multi-Principle Steering Deliberative Alignment: Reasoning Enables Safer Language Models

Reference 6

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T19:31:08.576351Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-08T11:49:47.994456Z digest=sha256:8975f04d2b22905e92e93f404ab7b28b6b6a24d15279b235806f55ba69e8557f

Observation 479779e5-6147-404b-84aa-3e672a7f678a · inbound

Internalizing Safety Understanding in Large Reasoning Models via Verification cites this paper.

Internalizing Safety Understanding in Large Reasoning Models via Verification Deliberative Alignment: Reasoning Enables Safer Language Models

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-12T01:51:14.318720Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-12T01:50:59.283409Z digest=sha256:55f75f42b444f1159950a124b483e2a7f60845dbe129c52be34ed07436beb3ac

Observation 272fc79d-4ef2-4ae5-ba07-222ec5920666 · inbound

Understanding Goal Generalisation in Sequential Reinforcement Learning cites this paper.

Understanding Goal Generalisation in Sequential Reinforcement Learning Deliberative Alignment: Reasoning Enables Safer Language Models

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-25T04:50:20.921471Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-25T04:49:50.034743Z digest=sha256:b96c83da545904c6f6221054d2e814a7d3d60c224194ef1bb9841fc85d3fa953

Observation 7f4d29f4-5a4d-4ba7-a7c7-0bbc2ddd70f1 · inbound

How Well Do Models Follow Their Constitutions? cites this paper.

How Well Do Models Follow Their Constitutions? Deliberative Alignment: Reasoning Enables Safer Language Models

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-06-30T15:34:47.874528Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-30T15:33:43.569710Z digest=sha256:28ac695366f34102219a5c08d83a3b4148ce46fdf16b4ed54a1cd2b4c78db345

Observation cdfbb531-120e-466a-b884-6941277f5414 · inbound

Reflect-Guard: Enhancing LLM Safeguards against Adversarial Prompts via Logical Self-Reflection cites this paper.

Reflect-Guard: Enhancing LLM Safeguards against Adversarial Prompts via Logical Self-Reflection Deliberative Alignment: Reasoning Enables Safer Language Models

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-07-01T16:35:50.610720Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-30T00:29:50.433221Z digest=sha256:db901b0706088f356fddcb6a97cfa697052b6c95c6cb7dcc9c890d217b61cb0a

Observation ba830e08-2ec1-429e-ba58-0d5e4ad252e6 · inbound

Mitigating Adaptive Attacks against Reasoning Models with Activation Consistency Training cites this paper.

Mitigating Adaptive Attacks against Reasoning Models with Activation Consistency Training Deliberative Alignment: Reasoning Enables Safer Language Models

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-06-29T14:23:30.706723Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-29T14:18:16.033063Z digest=sha256:4bf50797e17bf6a69e007787c065a0bf2d8d88dce96d5702371ad9e42afb0139

Observation 261b5b94-bf0f-4a3a-b4cb-96c188b31576 · inbound

Characterize Then Distill: Mechanistic Reasoning in Large Output Spaces cites this paper.

Characterize Then Distill: Mechanistic Reasoning in Large Output Spaces Deliberative Alignment: Reasoning Enables Safer Language Models

Reference 100

Resolution
verified exact
arxiv_id, observed 2026-06-27T22:31:21.526096Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-27T22:22:52.690010Z digest=sha256:b82d196381e369e5e3d8738da84893e1e5f5d19f9c8373e470f87925a7bb43cd

Observation 9967d366-91e4-4062-8624-8761f1e45cc6 · inbound

Activation Steering Induces Emergent Misalignment: A More Comprehensive Evaluation cites this paper.

Activation Steering Induces Emergent Misalignment: A More Comprehensive Evaluation Deliberative Alignment: Reasoning Enables Safer Language Models

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-07-02T22:57:26.124147Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-27T18:35:21.388956Z digest=sha256:37afd6d61eb536b6553e67ff698c893a1b8174da55526e81d15a40c0f9c3d7a8

Observation 0a3bd05e-3e5b-42b0-866e-0586952688a2 · inbound

Proxy Reward Internalization and Mechanistic Exploitation: A Learned Precursor to Reward Hacking and Its Generalization cites this paper.

Proxy Reward Internalization and Mechanistic Exploitation: A Learned Precursor to Reward Hacking and Its Generalization Deliberative Alignment: Reasoning Enables Safer Language Models

Reference 244

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T01:37:30.453603Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-27T16:26:34.918099Z digest=sha256:8095bed3b7213f3e03a28b31ce4bc8c3f2e83ece05387a0b85ea5c1fa7afc9b5

Observation bd46bb0c-007c-4699-a9ec-b322b81061fc · inbound

Greedy Coordinate Diffusion: Effective and Semantically Coherent Adversarial Attacks via Diffusion Guidance cites this paper.

Greedy Coordinate Diffusion: Effective and Semantically Coherent Adversarial Attacks via Diffusion Guidance Deliberative Alignment: Reasoning Enables Safer Language Models

Reference 17

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T17:08:43.430699Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-27T04:35:35.594085Z digest=sha256:2392d68bc66d69cb5bf2f8bac338243dff2983b4167e79d477eebde1e6744bce

Observation adb4d015-68ac-43f9-bbdb-4708714ccbe5 · inbound

Analyzing Defensive Misdirection Against Model-Guided Automated Attacks on Agentic AI Systems cites this paper.

Analyzing Defensive Misdirection Against Model-Guided Automated Attacks on Agentic AI Systems Deliberative Alignment: Reasoning Enables Safer Language Models

Reference 11

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T04:19:34.767260Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-26T17:04:38.863201Z digest=sha256:2fc9a0463d689550c5963cd17f19c136a28435c3793dd1f6eef16abbdbd31287

Observation 043c0c62-8bf8-42ae-b1ee-ea97acc3187d · inbound

Analyzing Defensive Misdirection Against Model-Guided Automated Attacks on Agentic AI Systems cites this paper.

Analyzing Defensive Misdirection Against Model-Guided Automated Attacks on Agentic AI Systems Deliberative Alignment: Reasoning Enables Safer Language Models

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-06-30T10:34:36.560267Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-30T10:30:03.047181Z digest=sha256:d18d71304eb5e9c3b1d9bcc79db847e219936055a9e09dfdbf26f18f9bf40356

Observation 125b2c5f-13ce-4795-8bb9-9e81e6ad67a1 · inbound

Do Thinking Tokens Help with Safety? cites this paper.

Do Thinking Tokens Help with Safety? Deliberative Alignment: Reasoning Enables Safer Language Models

Reference 7

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T17:30:00.639214Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-25T23:37:49.412578Z digest=sha256:3ca4db3d27df4b87384dc0868bec0caf38c67895b28fe1bec689136a8920bb0f

Observation 8e26d784-e1d1-4784-a47d-feedb03f575d · inbound

PolicyAlign: Direct Policy-Based Safety Alignment for Large Language Models cites this paper.

PolicyAlign: Direct Policy-Based Safety Alignment for Large Language Models Deliberative Alignment: Reasoning Enables Safer Language Models

Reference 17

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T19:40:06.722457Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-25T21:09:19.727723Z digest=sha256:7c782fff7a33483a3e7875fd9c932871aca44d20a296fe383965fc9e8d6c7f31

Observation 872dc00c-6b2b-4f47-8a70-0014c8a804a7 · inbound

Agent Safety Is Action Alignment cites this paper.

Agent Safety Is Action Alignment Deliberative Alignment: Reasoning Enables Safer Language Models

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-06-30T09:54:35.001200Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-30T09:50:45.759936Z digest=sha256:fe8fe7cf79cd606ba061b46c65a487a572e091924e30c62198c57c0dc971b4fa

Observation 1d0be5d4-c9cd-4fef-8161-6bc4d6f3d461 · inbound

HARC: Coupling Harmfulness and Refusal Directions for Robust Safety Alignment cites this paper.

HARC: Coupling Harmfulness and Refusal Directions for Robust Safety Alignment Deliberative Alignment: Reasoning Enables Safer Language Models

Reference 15

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T13:06:58.846398Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-02T12:57:00.805343Z digest=sha256:64cd0426b0aa3bb3d7b3847e38750878efa383796b56951cf1dcda2ad61818a5

Observation 9b2a054d-3b7f-4ee8-bccf-a3d47cd92f47 · inbound

HARC: Coupling Harmfulness and Refusal Directions for Robust Safety Alignment cites this paper.

HARC: Coupling Harmfulness and Refusal Directions for Robust Safety Alignment Deliberative Alignment: Reasoning Enables Safer Language Models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-07-12T09:27:01.450708Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T09:27:01.450708Z digest=sha256:5a02788bb1b4a9de7f60ee80500b4964dece5483cbdeea3189e617aa18b55fd4

Observation c34d92c6-d355-4da5-a9f0-5f14463cc273 · inbound

Cognitive Firewall: A Proactive, Zero-Trust, Multi-Gate Framework for LLM Safety cites this paper.

Cognitive Firewall: A Proactive, Zero-Trust, Multi-Gate Framework for LLM Safety Deliberative Alignment: Reasoning Enables Safer Language Models

Reference 15

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T20:38:54.934420Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-03T20:38:16.610310Z digest=sha256:e3fc5f8085f44d3858fd14419145c422eb4af6fc5cd38acef1eccb848cfe5fae

Observation b9561dc8-97ce-42f3-af4f-499781103678 · inbound

Oyster-II: Reinforcement Learning for Constructive Safety Alignment in Large Language Models cites this paper.

Oyster-II: Reinforcement Learning for Constructive Safety Alignment in Large Language Models Deliberative Alignment: Reasoning Enables Safer Language Models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-07-12T06:11:38.403281Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T06:11:38.403281Z digest=sha256:6cc4c2853fc57b98331f05f2b68d73b759bfc24d9f2880ca78a5bfb7c5a1fd23

Observation 5be75692-127e-4334-b317-c72f709d75c2 · inbound

Position: Preventing AI-Generated CSAM Necessitates New Approaches to AI Safety cites this paper.

Position: Preventing AI-Generated CSAM Necessitates New Approaches to AI Safety Deliberative Alignment: Reasoning Enables Safer Language Models

Reference 99

Resolution
unresolved
no resolver link, observed 2026-07-12T14:28:50.627444Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-12T14:28:50.627444Z digest=sha256:01b53704864238e7e8a0a4dabe96f88fd04943a571be77cfde29f3af2b28093b

Observation 7c8e5ee8-1db0-4ee8-b32e-b1651172bed9 · inbound

Cost of Reasoning in non-English Languages: A Case Study on Japanese cites this paper.

Cost of Reasoning in non-English Languages: A Case Study on Japanese Deliberative Alignment: Reasoning Enables Safer Language Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-07-14T14:12:44.286699Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T14:12:44.286699Z digest=sha256:8dd7edcf98ecae19c3d3b2df3a8eb17df119419d8305eef4fb6cc8aecce11098

Observation d6aed98b-cb00-497b-991a-2288d69fe821 · inbound

Underwriting the Agent Economy: The Blueprint for an AI Insurance Stack cites this paper.

Underwriting the Agent Economy: The Blueprint for an AI Insurance Stack Deliberative Alignment: Reasoning Enables Safer Language Models

Reference 230

Resolution
unresolved
no resolver link, observed 2026-07-15T08:35:47.870083Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-15T08:35:47.870083Z digest=sha256:c5120b6168bc0c50572a8693e2c02979475de821e97a0ac817c092509709cd27

Observation a01fcd39-8be5-4f88-84ba-1f363099a19e · inbound

Underwriting the Agent Economy: The Blueprint for an AI Insurance Stack cites this paper.

Underwriting the Agent Economy: The Blueprint for an AI Insurance Stack Deliberative Alignment: Reasoning Enables Safer Language Models

Reference 231

Resolution
unresolved
no resolver link, observed 2026-08-02T06:46:40.391719Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T06:46:40.391719Z digest=sha256:91324ad122352b6323df589c62d6d1e92ab467ed55a3018586f73c954622d5da

Observation d406d612-cdeb-4712-a465-bb94445022c3 · inbound

Verbalizable Representations Form a Global Workspace in Language Models cites this paper.

Verbalizable Representations Form a Global Workspace in Language Models Deliberative Alignment: Reasoning Enables Safer Language Models

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-01T23:15:23.274320Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:15:23.274320Z digest=sha256:c7be662778be94c2599d28e1ed441b71b55c841223baccdce5883e8220ab1276

Observation 4492457e-d1d8-4fe6-ac1c-0f87fa2c6a30 · inbound

A Geometric Perspective on Stabilizing Value Conflict Resolution cites this paper.

A Geometric Perspective on Stabilizing Value Conflict Resolution Deliberative Alignment: Reasoning Enables Safer Language Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-01T16:35:30.952522Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T16:35:30.952522Z digest=sha256:92917d133c1fa22dcd6866900ccb1e59d3a17e06146d5894f411e34e08cb44ff

Observation 79cea9df-6c32-4571-bcc4-7adcd1b3fb91 · inbound

QuantiBias: Benchmarking Quantization-Induced Bias in LLMs cites this paper.

QuantiBias: Benchmarking Quantization-Induced Bias in LLMs Deliberative Alignment: Reasoning Enables Safer Language Models

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-01T08:38:55.560417Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T08:38:55.560417Z digest=sha256:b7b75b67b0cf34c965764288d5e7c0609f2056482629c1537623a03787bce7e5

Observation d4c5bf99-e663-4e71-8425-5902ed995ed8 · inbound

Shared SFT Lessons Across Alignment, Model Organisms, and Toy Models cites this paper.

Shared SFT Lessons Across Alignment, Model Organisms, and Toy Models Deliberative Alignment: Reasoning Enables Safer Language Models

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-01T00:38:41.667721Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T00:38:41.667721Z digest=sha256:5c95c249465687d7c26e1b319d42ca7e37410f6745fffd3f5e50e9f388f1e583

Observation 78a191ed-6886-4149-8e7f-11aeed5beb12 · inbound

Constitutional Midtraining: Content Presence Drives Alignment Gains cites this paper.

Constitutional Midtraining: Content Presence Drives Alignment Gains Deliberative Alignment: Reasoning Enables Safer Language Models

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-01T11:35:00.225163Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T11:35:00.225163Z digest=sha256:c51617a9dacb69782d29dc7f9cd8015ec905adab5ba729aed441dd4902584b48

Observation db5790cb-4fa6-40dc-a93d-2b4355464f2f · inbound

Chain-of-Models: Cross-Model Auditing for Bias-Robust LLM Judges cites this paper.

Chain-of-Models: Cross-Model Auditing for Bias-Robust LLM Judges Deliberative Alignment: Reasoning Enables Safer Language Models

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-03T00:55:26.100966Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T00:55:26.100966Z digest=sha256:6b1ccb6ff8503d190fcb165e49360143cf121d447e3dc3bea3f94409823a4aa0

Observation 77fe7ada-8310-426f-8ddf-ff0fc104378d · inbound

Outcome-Guided Distillation: A Teacher-Student Framework to Advance VLM Reasoning in Autonomous Driving cites this paper.

Outcome-Guided Distillation: A Teacher-Student Framework to Advance VLM Reasoning in Autonomous Driving Deliberative Alignment: Reasoning Enables Safer Language Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-03T14:38:56.180832Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T14:38:56.180832Z digest=sha256:58d507af41687af391b43c723c2a017de0fd13138272d337df71cc33c9189480