Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-07-13T18:51:10.298187Z
Paper Citation Record · LEDGER
As of 7 August 2026, this Paper Citation Record lists 31 of 31 outbound references and 0 inbound Pith citation observations for arXiv:2606.00003.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-07-13T18:51:10.298187Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
31 of 31 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation bdc40f77-d8d7-474f-bb2e-03c83a3cd695 · outbound
Learning from Mistakes: Can LLM Self-Recover after Misalignment? Bengio, S
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 463068a0-b120-4749-8633-2895fb074331 · outbound
Learning from Mistakes: Can LLM Self-Recover after Misalignment? Unresolved cited work
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9981f573-b0ea-4bf6-98ae-a2fa9b8fa245 · outbound
Learning from Mistakes: Can LLM Self-Recover after Misalignment? Fine-Tuning Language Models from Human Preferences
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e249103c-ae54-42d1-82d1-4606e40ff225 · outbound
Learning from Mistakes: Can LLM Self-Recover after Misalignment? Learning to summarize from human feedback
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 37ed7e39-f930-4f14-b752-6a8ea0069950 · outbound
Learning from Mistakes: Can LLM Self-Recover after Misalignment? Training language models to follow instructions with human feedback
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 311df508-9ef0-4dfa-b278-44c36dd38923 · outbound
Learning from Mistakes: Can LLM Self-Recover after Misalignment? Direct Preference Optimization: Your Language Model is Secretly a Reward Model
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 44ca9a16-ffa9-4ac3-848b-b4efc6e34630 · outbound
Learning from Mistakes: Can LLM Self-Recover after Misalignment? Baseline Defenses for Adversarial Attacks Against Aligned Language Models
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5e6ea70a-49cb-425d-bb61-b5a8cff0c5f5 · outbound
Learning from Mistakes: Can LLM Self-Recover after Misalignment? Defending Against Alignment-Breaking Attacks via Robustly Aligned LLM
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 438d1711-9f9c-4492-a0d8-2e920a86953c · outbound
Learning from Mistakes: Can LLM Self-Recover after Misalignment? How Johnny Can Persuade LLMs to Jailbreak Them: Rethinking Persuasion to Challenge AI Safety by Humanizing LLMs
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7b58b01b-ec55-4260-9669-46bfa00262d1 · outbound
Learning from Mistakes: Can LLM Self-Recover after Misalignment? Unresolved cited work
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5c640ea5-1d00-4b98-a867-192691d3310c · outbound
Learning from Mistakes: Can LLM Self-Recover after Misalignment? Great, Now Write an Article About That: The Crescendo Multi-Turn LLM Jailbreak Attack
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 62cfdef6-7a62-43b0-bf7f-7c2f9be23ec6 · outbound
Learning from Mistakes: Can LLM Self-Recover after Misalignment? Universal and Transferable Adversarial Attacks on Aligned Language Models
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e28711a1-022a-48a8-919b-5e37e011090f · outbound
Learning from Mistakes: Can LLM Self-Recover after Misalignment? Towards Scalable Automated Alignment of LLMs: A Survey
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 00c58042-158f-4070-9cf8-bd4a1b5f5813 · outbound
Learning from Mistakes: Can LLM Self-Recover after Misalignment? Introducing v0.5 of the AI Safety Benchmark from MLCommons
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 084f4c77-a2d6-49cf-b82d-644caaf018a3 · outbound
Learning from Mistakes: Can LLM Self-Recover after Misalignment? Zheng, W.-L
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 32ae1346-80d1-4a3a-9782-f5207f2deb51 · outbound
Learning from Mistakes: Can LLM Self-Recover after Misalignment? Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation eb39685d-99f8-4feb-8372-d0c3f3da8c64 · outbound
Learning from Mistakes: Can LLM Self-Recover after Misalignment? G-Eval: NLG Evaluation using GPT-4 with Better Human Alignment
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 99714881-b8ce-445a-a233-d5c2f68a9d37 · outbound
Learning from Mistakes: Can LLM Self-Recover after Misalignment? Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a9363380-b9c8-4367-8308-1071695bacd0 · outbound
Learning from Mistakes: Can LLM Self-Recover after Misalignment? Jindal, H
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 38d4d4fe-8f22-471a-a6c7-bcee28f88cf5 · outbound
Learning from Mistakes: Can LLM Self-Recover after Misalignment? Unresolved cited work
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6db26b5f-812e-42eb-a53b-e31c082d7b5c · outbound
Learning from Mistakes: Can LLM Self-Recover after Misalignment? An Empirical Study of LLM-as-a-Judge: How Design Choices Impact Evaluation Reliability
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 95c790ca-7fc2-4ab8-9e9c-0e341194407e · outbound
Learning from Mistakes: Can LLM Self-Recover after Misalignment? A Holistic Approach to Undesired Content Detection in the Real World
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 655ccc02-a165-47fa-9d79-6c4bf50d95ca · outbound
Learning from Mistakes: Can LLM Self-Recover after Misalignment? Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cf62a882-9ef5-4ea2-a133-349da1568eb6 · outbound
Learning from Mistakes: Can LLM Self-Recover after Misalignment? Constitutional AI: Harmlessness from AI Feedback
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2fb57d64-3300-4884-9ed3-b83beb929b58 · outbound
Learning from Mistakes: Can LLM Self-Recover after Misalignment? Accessed: 2025
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 34a6731d-7ffc-462b-8994-6888630cd5de · outbound
Learning from Mistakes: Can LLM Self-Recover after Misalignment? Qwen3Guard Technical Report
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 59eda368-d3f5-41cf-b0ca-105bdbb18c47 · outbound
Learning from Mistakes: Can LLM Self-Recover after Misalignment? Unresolved cited work
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9946a1cd-f8e2-4ad4-82e5-f7b21607fdff · outbound
Learning from Mistakes: Can LLM Self-Recover after Misalignment? LLM Defenses Are Not Robust to Multi-Turn Human Jailbreaks Yet
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 16631ab4-7ae8-4856-bc12-1ba812be1c5a · outbound
Learning from Mistakes: Can LLM Self-Recover after Misalignment? Orlando, L
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fb39f707-e20d-455d-82bd-6fa876693a56 · outbound
Learning from Mistakes: Can LLM Self-Recover after Misalignment? Unresolved cited work
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 42628723-a748-4437-bd4e-ddb594436e86 · outbound
Learning from Mistakes: Can LLM Self-Recover after Misalignment? Unresolved cited work
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.