Pith. sign in

Paper Citation Record · LEDGER

Learning Safety Constraints for Large Language Models

As of 8 August 2026, this Paper Citation Record lists 25 of 25 outbound references and 5 inbound Pith citation observations for arXiv:2505.24445.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.24445 v1

Coverage vector

measured 25 of 25 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T12:30:16.628557Z

measured 30 of 30 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 5 of 5 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-02T11:57:51.928730Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

25 of 25 outbound references displayed

  • verified exact1
  • verified fuzzy2
  • unresolved20
  • parse uncertain0
  • malformed identifier2
  • metadata mismatch0

External citation measurements

0
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation 211c39db-1a9e-44e8-952e-e3a33be7cf9e · outbound

This paper cites MT-Bench-101: A Fine-Grained Benchmark for Evaluating Large Language Models in Multi-Turn Dialogues.

Learning Safety Constraints for Large Language Models MT-Bench-101: A Fine-Grained Benchmark for Evaluating Large Language Models in Multi-Turn Dialogues

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T12:30:14.985484Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:30:14.985484Z digest=sha256:70d5517de5bb4728c4b2f809f150803aacdd41eaa7bfec1fc3a9c4100d59643a

Observation 70a87e9a-0d49-4527-b423-512a812a02e6 · outbound

This paper cites SequenceMatch: Imitation Learning for Autoregressive Sequence Modelling with Backtracking.

Learning Safety Constraints for Large Language Models SequenceMatch: Imitation Learning for Autoregressive Sequence Modelling with Backtracking

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T12:30:15.354687Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:30:15.354687Z digest=sha256:cdba18c36e6bea80f18987ef0a5e0bb978fc132d81dee349857a0fc299e9a101

Observation 7363c4c8-40c6-482c-b8e0-7d7bc2524288 · outbound

This paper cites Safe Exploration in Continuous Action Spaces.

Learning Safety Constraints for Large Language Models Safe Exploration in Continuous Action Spaces

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T12:30:15.456333Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:30:15.456333Z digest=sha256:9d029ab7d9a9bd67def41fde9d7971b6d15088b7108d1e23965fcfbd0806ef84

Observation 08849323-40da-4714-a7fe-c4502c5a2c33 · outbound

This paper cites RealToxicityPrompts: Evaluating Neural Toxic Degeneration in Language Models.

Learning Safety Constraints for Large Language Models RealToxicityPrompts: Evaluating Neural Toxic Degeneration in Language Models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T12:30:15.620283Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:30:15.620283Z digest=sha256:82e232ba828b83f9692ac6219ca4636f1672b3441146a8dc187c87f320483721

Observation 519a09b1-99c1-4f8d-bf1e-1b428aa01b35 · outbound

This paper cites Measuring Massive Multitask Language Understanding.

Learning Safety Constraints for Large Language Models Measuring Massive Multitask Language Understanding

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T12:30:15.758569Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:30:15.758569Z digest=sha256:ad17e39afe4766418965289c61535bf40654e7c8cf63a4bd6369064ecb2556ee

Observation d78c2c74-df8f-49c9-87d0-4dd07e0a2639 · outbound

This paper cites Backdoor Attacks for In-Context Learning with Language Models.

Learning Safety Constraints for Large Language Models Backdoor Attacks for In-Context Learning with Language Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T12:30:15.821630Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:30:15.821630Z digest=sha256:790b04ea5c8fde59348a1490e71836251fe329e0358f3489aaadffdfbae84060

Observation 56cec1a0-9d3c-46d4-b90a-bb3b4cf916c4 · outbound

This paper cites AutoDAN: Generating Stealthy Jailbreak Prompts on Aligned Large Language Models.

Learning Safety Constraints for Large Language Models AutoDAN: Generating Stealthy Jailbreak Prompts on Aligned Large Language Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T12:30:15.939144Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:30:15.939144Z digest=sha256:cbde7acb441897973d0c9d6034c684cef40a0bdf1b5d6946c4102e7a204ce3de

Observation 7f1d5073-9f8a-45ce-be38-8a0bd2e8880e · outbound

This paper cites Enhancing LLM Safety via Constrained Direct Preference Optimization.

Learning Safety Constraints for Large Language Models Enhancing LLM Safety via Constrained Direct Preference Optimization

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T12:30:16.012101Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:30:16.012101Z digest=sha256:2240a91835279d72930f0b4fc40c2e870fcd1d78d629de780f73424dbff36dbc

Observation 41797f56-e347-427e-879c-d3aa95e037af · outbound

This paper cites Mpax: Mathematical pro- gramming in jax.

Learning Safety Constraints for Large Language Models Mpax: Mathematical pro- gramming in jax

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T12:30:16.069748Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:30:16.069748Z digest=sha256:ebc2c5bce4fb8d71280039d4cfe44cc4af0e6b77f5bb779fdf7fd70d4063f7dc

Observation 5df8c693-4082-40b4-82ae-d1d98aff73cb · outbound

This paper cites Enhancing Safety in Reinforcement Learning with Human Feedback via Rectified Policy Optimization.

Learning Safety Constraints for Large Language Models Enhancing Safety in Reinforcement Learning with Human Feedback via Rectified Policy Optimization

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T12:30:16.143593Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:30:16.143593Z digest=sha256:258c5776080cf45e2c3089ca5f6d7b5b71c9ee64f496cdb755a5e1cb7825d3ae

Observation 2481352e-529d-4429-8213-38d98983fd1c · outbound

This paper cites Constitutional Classifiers: Defending against Universal Jailbreaks across Thousands of Hours of Red Teaming.

Learning Safety Constraints for Large Language Models Constitutional Classifiers: Defending against Universal Jailbreaks across Thousands of Hours of Red Teaming

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T12:30:16.217942Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:30:16.217942Z digest=sha256:ba70f4da8c583a0e648e0b13513360940272e833912e9181dfa233f04d474552

Observation b459100a-c042-471c-be37-2187899fe70f · outbound

This paper cites Latent Adversarial Training Improves Robustness to Persistent Harmful Behaviors in LLMs.

Learning Safety Constraints for Large Language Models Latent Adversarial Training Improves Robustness to Persistent Harmful Behaviors in LLMs

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T12:30:16.287165Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:30:16.287165Z digest=sha256:ae055c0ac808e577855ee2f1c404a0c79782d8e70696de82fa6a05919e5c028b

Observation 74b81ec1-c15b-4fb3-a0bb-fb078d10d9bd · outbound

This paper cites Token-level Direct Preference Optimization.

Learning Safety Constraints for Large Language Models Token-level Direct Preference Optimization

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T12:30:16.341645Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:30:16.341645Z digest=sha256:d29215e07d4916bd2479cf1934787c65ba1d17dd1b3ed59a4c84bbfdc8a54ea2

Observation d0b25abb-0b16-4501-8622-7fd70047f051 · outbound

This paper cites Panacea: Pareto Alignment via Preference Adaptation for LLMs.

Learning Safety Constraints for Large Language Models Panacea: Pareto Alignment via Preference Adaptation for LLMs

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T12:30:16.406578Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:30:16.406578Z digest=sha256:0600b8dc9333bdf00e36c866581aa8cb8d9ecc7c9c54fce8172614c5bd8323fa

Observation cc3e2173-fffc-4d37-aa3e-989ffb38cc47 · outbound

This paper cites Beyond one-preference-fits-all alignment: Multi-objective direct preference optimization.

Learning Safety Constraints for Large Language Models Beyond one-preference-fits-all alignment: Multi-objective direct preference optimization

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:30:17.936203Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:30:16.462865Z digest=sha256:61694cd2a956bbbe1d5db74dc9d3dac68d6fdecb980cfa9a4eeb38d871639c09

Observation c06c74f7-cc17-4d49-8621-04f06d5d164c · outbound

This paper cites an unresolved cited work.

Learning Safety Constraints for Large Language Models Unresolved cited work

Reference 23

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:30:17.725886Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:30:16.506016Z digest=sha256:4712008ed7ffafb99999ed14f2cb161fab781f9311b01965060013a945e6b4cd

Observation 947a66c1-85a2-4efe-b390-8d096005b368 · outbound

This paper cites {human question}\n{model answer}.

Learning Safety Constraints for Large Language Models {human question}\n{model answer}

Reference 24

Resolution
malformed identifier
raw_fallback, observed 2026-08-07T12:30:17.579256Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:30:16.558669Z digest=sha256:0c4a41a1c67b7b1c546edc5f1ae56671725ee04969d02e5311afa02e22d8ac5b

Observation e2f0e159-4e83-4861-9165-53506c0bd2b9 · outbound

This paper cites Results show mean ± standard deviation.

Learning Safety Constraints for Large Language Models Results show mean ± standard deviation

Reference 25

Resolution
malformed identifier
raw_fallback, observed 2026-08-07T12:30:17.301765Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:30:16.628557Z digest=sha256:3150590447074288426c658a30764f24a9f0be35ed9373a3ba2c29f64acf0002

Observation 12ed9f93-7ac7-444e-845a-2394f397f838 · outbound

This paper cites Jailbreaking Leading Safety-Aligned LLMs with Simple Adaptive Attacks.

Learning Safety Constraints for Large Language Models Jailbreaking Leading Safety-Aligned LLMs with Simple Adaptive Attacks

Reference 1999

Resolution
unresolved
no resolver link, observed 2026-08-07T12:30:14.786704Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:30:14.786704Z digest=sha256:9c281473fe31016e5d43459abed8730d279c4af546c35c84794ff59402e21c7a

Observation 3fc8dcb9-694d-40f7-bfa6-435116fa0bca · outbound

This paper cites TruthfulQA: Measuring How Models Mimic Human Falsehoods.

Learning Safety Constraints for Large Language Models TruthfulQA: Measuring How Models Mimic Human Falsehoods

Reference 2014

Resolution
unresolved
no resolver link, observed 2026-08-07T12:30:15.888021Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:30:15.888021Z digest=sha256:299327cea3099bada3305dd1f3753b454fdf2a3ab747303970ca9fa68cb63f99

Observation 5691e74e-0f86-4b94-8d91-7297b889fcc0 · outbound

This paper cites Interpreting Neural Networks through the Polytope Lens.

Learning Safety Constraints for Large Language Models Interpreting Neural Networks through the Polytope Lens

Reference 2016

Resolution
unresolved
no resolver link, observed 2026-08-07T12:30:15.138252Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:30:15.138252Z digest=sha256:e63bbda13a95bf2fb527eb7a91280dd8322b9c99cfaac82534f20e6b7145c6f8

Observation 62406f3f-8d92-4d13-9b40-ac547bef0427 · outbound

This paper cites AI Control: Improving Safety Despite Intentional Subversion.

Learning Safety Constraints for Large Language Models AI Control: Improving Safety Despite Intentional Subversion

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-07T12:30:15.698528Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:30:15.698528Z digest=sha256:47152050c4f7cc9861f37624d76181e999bbcaab479a89473dd71b66dace9365

Observation 4c4f088e-3a56-4d72-9850-0fc309a42dff · outbound

This paper cites Unlocking Decoding-time Controllability: Gradient-Free Multi-Objective Alignment with Contrastive Prompts.

Learning Safety Constraints for Large Language Models Unlocking Decoding-time Controllability: Gradient-Free Multi-Objective Alignment with Contrastive Prompts

Reference 2022

Resolution
verified exact
local_arxiv, observed 2026-08-07T12:30:17.032993Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:30:15.545665Z digest=sha256:d6302716e7f314ce17eb4e5c993b415d9479cd54328164f4b61cc5a23ee0bf1a

Observation 28a20280-8c32-43e8-b0aa-224937fb62cc · outbound

This paper cites Defending Against Unforeseen Failure Modes with Latent Adversarial Training.

Learning Safety Constraints for Large Language Models Defending Against Unforeseen Failure Modes with Latent Adversarial Training

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-07T12:30:15.271464Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:30:15.271464Z digest=sha256:15e434bdef788eb1aa466812761163e16c3c81ef833fb763f96e7ae76b4cf9a2

Observation 05e6391c-18c6-4105-a685-75cfa16f0747 · outbound

This paper cites and Bartlett, P.

Learning Safety Constraints for Large Language Models and Bartlett, P

Reference 2024

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:30:18.150799Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:30:14.835783Z digest=sha256:dc09f4840f5cf4a394c7f4f5dc93db4399d4b8b16b2a95c5c4d6bec2a7afc7bc

Pith citing papers

Observation e2bd960a-5dbb-4c95-9548-c457ff5dc172 · inbound

When control meets large language models: From words to dynamics cites this paper.

When control meets large language models: From words to dynamics Learning Safety Constraints for Large Language Models

Reference 257

Resolution
verified exact
arxiv_id, observed 2026-05-21T14:54:13.126908Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-21T14:52:44.632671Z digest=sha256:9d7a77f7121509ac37cab8ace46bf39aa26b0b6eb871dd7dea7f05745817ab7b

Observation 9036905c-b097-4870-b700-ba09c4020dc2 · inbound

Preference Instability in Reward Models: Detection and Mitigation via Sparse Autoencoders cites this paper.

Preference Instability in Reward Models: Detection and Mitigation via Sparse Autoencoders Learning Safety Constraints for Large Language Models

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-20T22:49:10.216730Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-20T22:48:54.238767Z digest=sha256:eecdce237558ded8f461d9f2309d1db6b4d850cfee8625a2fc86518f7a3623e9

Observation c9ba4192-0e7a-4ee5-ab2c-d234dde2bced · inbound

Defending Jailbreak Attacks on Large Language Models via Manifold Trajectory Kinetics cites this paper.

Defending Jailbreak Attacks on Large Language Models via Manifold Trajectory Kinetics Learning Safety Constraints for Large Language Models

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-06-27T22:01:20.847532Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-27T21:55:48.561400Z digest=sha256:0a8ec32ad67003041c4d48af46d2e024ed4d6ddf3c9a2205d8b6824cae717495

Observation 511e8842-49c9-4e3a-a11b-475e12eb343f · inbound

Greedy Coordinate Diffusion: Effective and Semantically Coherent Adversarial Attacks via Diffusion Guidance cites this paper.

Greedy Coordinate Diffusion: Effective and Semantically Coherent Adversarial Attacks via Diffusion Guidance Learning Safety Constraints for Large Language Models

Reference 15

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T17:08:43.438995Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-27T04:35:35.594085Z digest=sha256:9bf7b0adc038e143524a9805f066c65eac8482175573c9465e73fea1d85f3c39

Observation 9d400a3f-0fe4-42db-8206-794618102f72 · inbound

Geometry-Guided Constraint Learning for LLM Safety Classification cites this paper.

Geometry-Guided Constraint Learning for LLM Safety Classification Learning Safety Constraints for Large Language Models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-02T11:57:51.928730Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T11:57:51.928730Z digest=sha256:3eefde361366e962ce4fe5c7650d21a6f267b337d11980288667ba8890c3518d