Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T12:30:16.628557Z
Paper Citation Record · LEDGER
As of 8 August 2026, this Paper Citation Record lists 25 of 25 outbound references and 5 inbound Pith citation observations for arXiv:2505.24445.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T12:30:16.628557Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-02T11:57:51.928730Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z
25 of 25 outbound references displayed
External citation measurements
0
arxiv_reference, observed 2026-08-05T02:28:24.338817Z
Observation 211c39db-1a9e-44e8-952e-e3a33be7cf9e · outbound
Learning Safety Constraints for Large Language Models MT-Bench-101: A Fine-Grained Benchmark for Evaluating Large Language Models in Multi-Turn Dialogues
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 70a87e9a-0d49-4527-b423-512a812a02e6 · outbound
Learning Safety Constraints for Large Language Models SequenceMatch: Imitation Learning for Autoregressive Sequence Modelling with Backtracking
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7363c4c8-40c6-482c-b8e0-7d7bc2524288 · outbound
Learning Safety Constraints for Large Language Models Safe Exploration in Continuous Action Spaces
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 08849323-40da-4714-a7fe-c4502c5a2c33 · outbound
Learning Safety Constraints for Large Language Models RealToxicityPrompts: Evaluating Neural Toxic Degeneration in Language Models
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 519a09b1-99c1-4f8d-bf1e-1b428aa01b35 · outbound
Learning Safety Constraints for Large Language Models Measuring Massive Multitask Language Understanding
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d78c2c74-df8f-49c9-87d0-4dd07e0a2639 · outbound
Learning Safety Constraints for Large Language Models Backdoor Attacks for In-Context Learning with Language Models
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 56cec1a0-9d3c-46d4-b90a-bb3b4cf916c4 · outbound
Learning Safety Constraints for Large Language Models AutoDAN: Generating Stealthy Jailbreak Prompts on Aligned Large Language Models
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7f1d5073-9f8a-45ce-be38-8a0bd2e8880e · outbound
Learning Safety Constraints for Large Language Models Enhancing LLM Safety via Constrained Direct Preference Optimization
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 41797f56-e347-427e-879c-d3aa95e037af · outbound
Learning Safety Constraints for Large Language Models Mpax: Mathematical pro- gramming in jax
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5df8c693-4082-40b4-82ae-d1d98aff73cb · outbound
Learning Safety Constraints for Large Language Models Enhancing Safety in Reinforcement Learning with Human Feedback via Rectified Policy Optimization
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2481352e-529d-4429-8213-38d98983fd1c · outbound
Learning Safety Constraints for Large Language Models Constitutional Classifiers: Defending against Universal Jailbreaks across Thousands of Hours of Red Teaming
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b459100a-c042-471c-be37-2187899fe70f · outbound
Learning Safety Constraints for Large Language Models Latent Adversarial Training Improves Robustness to Persistent Harmful Behaviors in LLMs
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 74b81ec1-c15b-4fb3-a0bb-fb078d10d9bd · outbound
Learning Safety Constraints for Large Language Models Token-level Direct Preference Optimization
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d0b25abb-0b16-4501-8622-7fd70047f051 · outbound
Learning Safety Constraints for Large Language Models Panacea: Pareto Alignment via Preference Adaptation for LLMs
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cc3e2173-fffc-4d37-aa3e-989ffb38cc47 · outbound
Learning Safety Constraints for Large Language Models Beyond one-preference-fits-all alignment: Multi-objective direct preference optimization
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation c06c74f7-cc17-4d49-8621-04f06d5d164c · outbound
Learning Safety Constraints for Large Language Models Unresolved cited work
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 947a66c1-85a2-4efe-b390-8d096005b368 · outbound
Learning Safety Constraints for Large Language Models {human question}\n{model answer}
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation e2f0e159-4e83-4861-9165-53506c0bd2b9 · outbound
Learning Safety Constraints for Large Language Models Results show mean ± standard deviation
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 12ed9f93-7ac7-444e-845a-2394f397f838 · outbound
Learning Safety Constraints for Large Language Models Jailbreaking Leading Safety-Aligned LLMs with Simple Adaptive Attacks
Reference 1999
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3fc8dcb9-694d-40f7-bfa6-435116fa0bca · outbound
Learning Safety Constraints for Large Language Models TruthfulQA: Measuring How Models Mimic Human Falsehoods
Reference 2014
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5691e74e-0f86-4b94-8d91-7297b889fcc0 · outbound
Learning Safety Constraints for Large Language Models Interpreting Neural Networks through the Polytope Lens
Reference 2016
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 62406f3f-8d92-4d13-9b40-ac547bef0427 · outbound
Learning Safety Constraints for Large Language Models AI Control: Improving Safety Despite Intentional Subversion
Reference 2021
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4c4f088e-3a56-4d72-9850-0fc309a42dff · outbound
Learning Safety Constraints for Large Language Models Unlocking Decoding-time Controllability: Gradient-Free Multi-Objective Alignment with Contrastive Prompts
Reference 2022
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 28a20280-8c32-43e8-b0aa-224937fb62cc · outbound
Learning Safety Constraints for Large Language Models Defending Against Unforeseen Failure Modes with Latent Adversarial Training
Reference 2023
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 05e6391c-18c6-4105-a685-75cfa16f0747 · outbound
Learning Safety Constraints for Large Language Models and Bartlett, P
Reference 2024
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation e2bd960a-5dbb-4c95-9548-c457ff5dc172 · inbound
When control meets large language models: From words to dynamics Learning Safety Constraints for Large Language Models
Reference 257
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 9036905c-b097-4870-b700-ba09c4020dc2 · inbound
Preference Instability in Reward Models: Detection and Mitigation via Sparse Autoencoders Learning Safety Constraints for Large Language Models
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation c9ba4192-0e7a-4ee5-ab2c-d234dde2bced · inbound
Defending Jailbreak Attacks on Large Language Models via Manifold Trajectory Kinetics Learning Safety Constraints for Large Language Models
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 511e8842-49c9-4e3a-a11b-475e12eb343f · inbound
Greedy Coordinate Diffusion: Effective and Semantically Coherent Adversarial Attacks via Diffusion Guidance Learning Safety Constraints for Large Language Models
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 9d400a3f-0fe4-42db-8206-794618102f72 · inbound
Geometry-Guided Constraint Learning for LLM Safety Classification Learning Safety Constraints for Large Language Models
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.