Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T15:48:01.488303Z
Paper Citation Record · LEDGER
As of 9 August 2026, this Paper Citation Record lists 54 of 54 outbound references and 3 inbound Pith citation observations for arXiv:2507.14987.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T15:48:01.488303Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-05T23:46:20.547519Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-04T19:40:06.671432Z
54 of 54 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 88ba8c8e-c977-4d46-b786-55e5cba25ab7 · outbound
AlphaAlign: Incentivizing Safety Alignment with Extremely Simplified Reinforcement Learning Claude 3.5 sonnet model card addendum
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation d80e5d73-1e38-4a03-81a6-2d3663420e85 · outbound
AlphaAlign: Incentivizing Safety Alignment with Extremely Simplified Reinforcement Learning Claude 3.7 sonnet system card
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 0bd55893-82c6-4e7f-b58f-c7a20b20f6f0 · outbound
AlphaAlign: Incentivizing Safety Alignment with Extremely Simplified Reinforcement Learning Gemma 2: Improving Open Language Models at a Practical Size
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3fe5fa3a-47f8-40a7-a1fb-c09fc21ac8d2 · outbound
AlphaAlign: Incentivizing Safety Alignment with Extremely Simplified Reinforcement Learning Qwen2.5 Technical Report
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b7dbd5c6-5592-49c6-ab0c-908c54cd6b8a · outbound
AlphaAlign: Incentivizing Safety Alignment with Extremely Simplified Reinforcement Learning DeepSeek-V3 Technical Report
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3169a0e9-52a7-4e5f-812e-830775b113a0 · outbound
AlphaAlign: Incentivizing Safety Alignment with Extremely Simplified Reinforcement Learning Hadi Amini, and Yanzhao Wu
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation e6826717-384c-4b13-b4d2-b0eedd7eb3d9 · outbound
AlphaAlign: Incentivizing Safety Alignment with Extremely Simplified Reinforcement Learning Trustworthy LLMs: a Survey and Guideline for Evaluating Large Language Models' Alignment
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c4445efa-845a-47c7-91c3-2a9bb63aef44 · outbound
AlphaAlign: Incentivizing Safety Alignment with Extremely Simplified Reinforcement Learning Safe + Safe = Unsafe? Exploring How Safe Images Can Be Exploited to Jailbreak Large Vision-Language Models
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2ce16309-30d6-4c2e-81ab-dbe6b378be84 · outbound
AlphaAlign: Incentivizing Safety Alignment with Extremely Simplified Reinforcement Learning Kummerfeld, and Rada Mihalcea
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 9e2ec079-d332-444a-9c5b-7f9aabb0ce09 · outbound
AlphaAlign: Incentivizing Safety Alignment with Extremely Simplified Reinforcement Learning Towards understanding jailbreak attacks in llms: A representation space analysis
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation ae7756c7-859a-47ee-9ec7-d1d2f800487f · outbound
AlphaAlign: Incentivizing Safety Alignment with Extremely Simplified Reinforcement Learning Uncovering safety risks of large language models through concept activation vector
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation b86170f1-ddd6-431b-80c1-1690a8b0d795 · outbound
AlphaAlign: Incentivizing Safety Alignment with Extremely Simplified Reinforcement Learning On prompt-driven safeguarding for large language models
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation f572c8e5-e434-4ba6-bbb2-951b37b78fb2 · outbound
AlphaAlign: Incentivizing Safety Alignment with Extremely Simplified Reinforcement Learning Nguyen, Jun Sun, and Tat - Seng Chua
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 337aa62b-04b5-42e8-a50f-5fb9a4304c0a · outbound
AlphaAlign: Incentivizing Safety Alignment with Extremely Simplified Reinforcement Learning Jailbroken: How does LLM safety training fail? In NeurIPS, 2023
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 0064e74f-d67b-4e40-b717-2c99ddc7614c · outbound
AlphaAlign: Incentivizing Safety Alignment with Extremely Simplified Reinforcement Learning Safety Reasoning with Guidelines
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 03d9e388-4c7d-45ee-9429-c8443e144e39 · outbound
AlphaAlign: Incentivizing Safety Alignment with Extremely Simplified Reinforcement Learning Rule based rewards for language model safety
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 1865f620-a909-40b7-a460-e45446f6d8f0 · outbound
AlphaAlign: Incentivizing Safety Alignment with Extremely Simplified Reinforcement Learning Safety is not only about refusal: Reasoning-enhanced fine-tuning for interpretable LLM safety
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a6aa999a-65fe-4a16-a1af-f85c3a9b6d6d · outbound
AlphaAlign: Incentivizing Safety Alignment with Extremely Simplified Reinforcement Learning Zhao, Kelvin Guu, Adams Wei Yu, Brian Lester, Nan Du, Andrew M
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 808083e7-b3a3-4e0c-ae05-a92ca6e7c8ac · outbound
AlphaAlign: Incentivizing Safety Alignment with Extremely Simplified Reinforcement Learning Unresolved cited work
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 259e10b5-eb2d-4174-be50-c6a1a9969e59 · outbound
AlphaAlign: Incentivizing Safety Alignment with Extremely Simplified Reinforcement Learning Manning, Stefano Ermon, and Chelsea Finn
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b502d8ad-ff9b-4891-9fb4-fa4f7923d7b2 · outbound
AlphaAlign: Incentivizing Safety Alignment with Extremely Simplified Reinforcement Learning Safe lora: The silver lining of reducing safety risks when finetuning large language models
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation c8dd559e-859e-45f9-8052-95d8bd935554 · outbound
AlphaAlign: Incentivizing Safety Alignment with Extremely Simplified Reinforcement Learning Zico Kolter, Matt Fredrikson, and Dan Hendrycks
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation c4b5728f-b8ff-4436-83de-53d40cfe133f · outbound
AlphaAlign: Incentivizing Safety Alignment with Extremely Simplified Reinforcement Learning Safety Alignment Should Be Made More Than Just a Few Tokens Deep
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 449e52f3-1076-4415-b5b0-64a3b80315c0 · outbound
AlphaAlign: Incentivizing Safety Alignment with Extremely Simplified Reinforcement Learning Deliberative Alignment: Reasoning Enables Safer Language Models
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1600e39f-f572-4f6c-a4ea-866b9f0be95d · outbound
AlphaAlign: Incentivizing Safety Alignment with Extremely Simplified Reinforcement Learning Enhancing model defense against jailbreaks with proactive safety reasoning
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c04d912f-6ad0-46de-9970-157ed05c112d · outbound
AlphaAlign: Incentivizing Safety Alignment with Extremely Simplified Reinforcement Learning Reasoning-to-defend: Safety-aware reasoning can defend large language models from jailbreaking
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6c62134f-a5b8-4b75-8a62-20223ca065e8 · outbound
AlphaAlign: Incentivizing Safety Alignment with Extremely Simplified Reinforcement Learning Bikel, Jason E
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation d50ed610-393a-461e-9e5c-49de9b04a66c · outbound
AlphaAlign: Incentivizing Safety Alignment with Extremely Simplified Reinforcement Learning Does Refusal Training in LLMs Generalize to the Past Tense?
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2b78b9f0-7a0e-4039-8e93-d406d8f2ae0c · outbound
AlphaAlign: Incentivizing Safety Alignment with Extremely Simplified Reinforcement Learning Fine-tuning aligned language models compromises safety, even when users do not intend to! In ICLR , 2024 b
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation f1929073-70ea-4d28-aba7-49977dfbb656 · outbound
AlphaAlign: Incentivizing Safety Alignment with Extremely Simplified Reinforcement Learning Constitutional AI: Harmlessness from AI Feedback
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 899225ce-cc30-48db-baea-f4a0d0aee895 · outbound
AlphaAlign: Incentivizing Safety Alignment with Extremely Simplified Reinforcement Learning DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 360512d6-0b7c-4d22-9ad5-1df3e4542611 · outbound
AlphaAlign: Incentivizing Safety Alignment with Extremely Simplified Reinforcement Learning Open-Reasoner-Zero: An Open Source Approach to Scaling Up Reinforcement Learning on the Base Model
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 81b675cb-451b-41e5-8199-e56ff022df83 · outbound
AlphaAlign: Incentivizing Safety Alignment with Extremely Simplified Reinforcement Learning Does Reinforcement Learning Really Incentivize Reasoning Capacity in LLMs Beyond the Base Model?
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d2109089-99c2-41ee-9788-e561901be489 · outbound
AlphaAlign: Incentivizing Safety Alignment with Extremely Simplified Reinforcement Learning Cognitive Behaviors that Enable Self-Improving Reasoners, or, Four Habits of Highly Effective STaRs
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9af60e70-f02b-4ebb-8450-da15badc50b0 · outbound
AlphaAlign: Incentivizing Safety Alignment with Extremely Simplified Reinforcement Learning DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5a80a100-a6ef-4799-9694-013717548600 · outbound
AlphaAlign: Incentivizing Safety Alignment with Extremely Simplified Reinforcement Learning Proximal Policy Optimization Algorithms
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a5ed7fa2-05ae-42b8-83bd-5e6dc2241560 · outbound
AlphaAlign: Incentivizing Safety Alignment with Extremely Simplified Reinforcement Learning Kimi k1.5: Scaling Reinforcement Learning with LLMs
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b6d57317-baef-4e62-ac3f-bf8c51e592c0 · outbound
AlphaAlign: Incentivizing Safety Alignment with Extremely Simplified Reinforcement Learning Jordan, and Pieter Abbeel
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 0f6e7206-c031-46f3-8e11-5fbed43cc04d · outbound
AlphaAlign: Incentivizing Safety Alignment with Extremely Simplified Reinforcement Learning A strongreject for empty jailbreaks
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 05a11772-9b8d-4347-87ad-2ea278f2011e · outbound
AlphaAlign: Incentivizing Safety Alignment with Extremely Simplified Reinforcement Learning Universal and Transferable Adversarial Attacks on Aligned Language Models
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cda7c6b2-92cc-4789-890d-0239ae4548ac · outbound
AlphaAlign: Incentivizing Safety Alignment with Extremely Simplified Reinforcement Learning Wildguard: Open one-stop moderation tools for safety risks, jailbreaks, and refusals of llms
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 04cb8820-72d2-4536-b12f-2c765052765e · outbound
AlphaAlign: Incentivizing Safety Alignment with Extremely Simplified Reinforcement Learning Unresolved cited work
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 53376565-e138-4c82-bfd6-a4b69299a99a · outbound
AlphaAlign: Incentivizing Safety Alignment with Extremely Simplified Reinforcement Learning Jailbreaking Black Box Large Language Models in Twenty Queries
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2de32363-6a72-45f3-a94f-7fa96500ef52 · outbound
AlphaAlign: Incentivizing Safety Alignment with Extremely Simplified Reinforcement Learning Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ddb9a189-4774-4efc-af9c-4a844b2ae6c8 · outbound
AlphaAlign: Incentivizing Safety Alignment with Extremely Simplified Reinforcement Learning Smith, Yejin Choi, and Hanna Hajishirzi
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 6a0dec33-4c25-4381-903b-8a2bfec30c1e · outbound
AlphaAlign: Incentivizing Safety Alignment with Extremely Simplified Reinforcement Learning Measuring massive multitask language understanding
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 9d460353-3ff9-41c4-b0c3-679149ea8655 · outbound
AlphaAlign: Incentivizing Safety Alignment with Extremely Simplified Reinforcement Learning Length-Controlled AlpacaEval: A Simple Way to Debias Automatic Evaluators
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 95f9e80f-c966-49b4-9e9e-1dc38a4d5fe4 · outbound
AlphaAlign: Incentivizing Safety Alignment with Extremely Simplified Reinforcement Learning Training Verifiers to Solve Math Word Problems
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ffa41b95-17d9-4168-811e-97642a8be57d · outbound
AlphaAlign: Incentivizing Safety Alignment with Extremely Simplified Reinforcement Learning The Llama 3 Herd of Models
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a9b04422-6b54-44f1-9b15-d7d196de2a5f · outbound
AlphaAlign: Incentivizing Safety Alignment with Extremely Simplified Reinforcement Learning Llama 2: Open Foundation and Fine-Tuned Chat Models
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 572296ea-4dd0-4882-9ef0-3338dad7fd00 · outbound
AlphaAlign: Incentivizing Safety Alignment with Extremely Simplified Reinforcement Learning Free dolly: Introducing the world's first truly open instruction-tuned llm, 2023
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7feb1d53-edbb-42e0-9f61-9bff21590fc5 · outbound
AlphaAlign: Incentivizing Safety Alignment with Extremely Simplified Reinforcement Learning Xstest: A test suite for identifying exaggerated safety behaviours in large language models
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 22bfb320-4556-4c0c-8a4a-51cafd65d2cc · outbound
AlphaAlign: Incentivizing Safety Alignment with Extremely Simplified Reinforcement Learning HybridFlow: A Flexible and Efficient RLHF Framework
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5c2cfbf5-5502-4743-bb2f-114e7d7fc1a7 · outbound
AlphaAlign: Incentivizing Safety Alignment with Extremely Simplified Reinforcement Learning Iterative preference learning from human feedback: Bridging theory and practice for rlhf under kl-constraint, 2024
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0ad39314-b1f0-485a-8b5f-70fe4ef8b5c0 · inbound
Self-ReSET: Learning to Self-Recover from Unsafe Reasoning Trajectories AlphaAlign: Incentivizing Safety Alignment with Extremely Simplified Reinforcement Learning
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation a1575c0b-db97-42d4-a424-7ccff2530cb6 · inbound
PolicyAlign: Direct Policy-Based Safety Alignment for Large Language Models AlphaAlign: Incentivizing Safety Alignment with Extremely Simplified Reinforcement Learning
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 83c5b1b7-a17d-416c-9347-9ab31cd1660e · inbound
When Refusal Looks Safe: The Refusal-Cue Shortcut in Safety Guard Models AlphaAlign: Incentivizing Safety Alignment with Extremely Simplified Reinforcement Learning
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.