Pith. sign in

Paper Citation Record · LEDGER

On Almost Surely Safe Alignment of Large Language Models at Inference-Time

As of 18 August 2026, this Paper Citation Record lists 100 of 125 outbound references and 2 inbound Pith citation observations for arXiv:2502.01208.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2502.01208 v3

Coverage vector

measured 100 of 125 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-09T16:18:40.969178Z

measured 102 of 102 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T12:02:43.148072Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-07T05:43:07.622958Z

Reference resolution

100 of 125 outbound references displayed

  • verified exact6
  • verified fuzzy0
  • unresolved94
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 8ac5fa90-7d16-4348-a5ad-6f20f83efd4c · outbound

This paper cites An empirical survey on long document summarization: Datasets, models, and metrics.ACM computing surveys, 55(8):1–35, 2022.

On Almost Surely Safe Alignment of Large Language Models at Inference-Time An empirical survey on long document summarization: Datasets, models, and metrics.ACM computing surveys, 55(8):1–35, 2022

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-09T16:18:40.470256Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T16:18:40.470256Z digest=sha256:64b5b3d10a66f2767b7c5136538ce516b6a7ab3dae3f716254277929570056e6

Observation 1181761a-d634-4615-af8e-de209e60ba4f · outbound

This paper cites Learning to summarize with human feedback.

On Almost Surely Safe Alignment of Large Language Models at Inference-Time Learning to summarize with human feedback

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-09T16:18:40.476598Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T16:18:40.476598Z digest=sha256:f4a10706b646e42956fef02c85accedb9ace69c610a94eb22f9a54cb283824c5

Observation b1e88c8a-2a73-44df-9c95-7a9713802ee6 · outbound

This paper cites Pal: Program-aided language models.

On Almost Surely Safe Alignment of Large Language Models at Inference-Time Pal: Program-aided language models

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-09T16:18:40.481857Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T16:18:40.481857Z digest=sha256:fe51c5bb75afefd0b5f8ec2a7626a94e322c31bb313fe246d0c13bfed085848a

Observation f6d39b8d-bfc0-4123-823f-368013e915a5 · outbound

This paper cites Evaluating Large Language Models Trained on Code.

On Almost Surely Safe Alignment of Large Language Models at Inference-Time Evaluating Large Language Models Trained on Code

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-09T16:18:40.487014Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T16:18:40.487014Z digest=sha256:a60bea25ea904f5c0d2d133ce33a81990902167c3f3613175cffb31618a0af81

Observation 1e67b72c-ab4c-4c90-9378-ac0cbdc4abfc · outbound

This paper cites ROS-LLM: A ROS framework for embodied AI with task feedback and structured reasoning.

On Almost Surely Safe Alignment of Large Language Models at Inference-Time ROS-LLM: A ROS framework for embodied AI with task feedback and structured reasoning

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-09T16:18:40.492246Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T16:18:40.492246Z digest=sha256:6aa9c470be97c748adde400b151e3624a9fb780fa348de62cdf16f0291d9215d

Observation af3910d5-6cb5-4ca6-8fb1-c8032dc86753 · outbound

This paper cites A sur- vey on integration of large language models with intelligent robots.Intelligent Service Robotics, 17(5):1091–1107, August 2024.

On Almost Surely Safe Alignment of Large Language Models at Inference-Time A sur- vey on integration of large language models with intelligent robots.Intelligent Service Robotics, 17(5):1091–1107, August 2024

Reference 6

Resolution
verified exact
doi, observed 2026-08-09T16:18:41.128413Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-09T16:18:40.497065Z digest=sha256:0ec123108b23320e5b92fc87265cb791a81754c4724c1ad52b46451aed011d29

Observation 2025b726-2696-4717-be04-c3c9820a338c · outbound

This paper cites Toxicity in ChatGPT: Analyzing Persona-assigned Language Models.

On Almost Surely Safe Alignment of Large Language Models at Inference-Time Toxicity in ChatGPT: Analyzing Persona-assigned Language Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-09T16:18:40.502533Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T16:18:40.502533Z digest=sha256:b5dd0b410033f0e7e4ebc77a74982ae9a715050197d9ff7e68eedc29b27cfa2b

Observation ef519bac-1552-4b43-ac33-199be15ae185 · outbound

This paper cites Red Teaming Language Models to Reduce Harms: Methods, Scaling Behaviors, and Lessons Learned.

On Almost Surely Safe Alignment of Large Language Models at Inference-Time Red Teaming Language Models to Reduce Harms: Methods, Scaling Behaviors, and Lessons Learned

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-09T16:18:40.507237Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T16:18:40.507237Z digest=sha256:654ab0a2395644eda7405fba4be3dcaee4456ef6cab08247007636d654bb8ecf

Observation 64a6923b-99a6-47c1-ab70-77ddc7620e35 · outbound

This paper cites Ethical and social risks of harm from Language Models.

On Almost Surely Safe Alignment of Large Language Models at Inference-Time Ethical and social risks of harm from Language Models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-09T16:18:40.512328Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T16:18:40.512328Z digest=sha256:34e1d19413ddc4a6731995a5811673ec3bc57c184f098397307657eb2c84fd58

Observation db4b25dc-92a5-41ff-b2e2-dd61d1ec7116 · outbound

This paper cites RealToxicityPrompts: Evaluating Neural Toxic Degeneration in Language Models.

On Almost Surely Safe Alignment of Large Language Models at Inference-Time RealToxicityPrompts: Evaluating Neural Toxic Degeneration in Language Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-09T16:18:40.517461Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T16:18:40.517461Z digest=sha256:4858b468d5e60c47a4e5608f75556a4260428fa085c52e73b4c8abcdebda3288

Observation f199c6c5-b8ec-4f5b-b0e0-a29eb46ccddb · outbound

This paper cites Training language models to follow instructions with human feedback.Advances in Neural Information Processing Systems, 35:27730–27744, 2022.

On Almost Surely Safe Alignment of Large Language Models at Inference-Time Training language models to follow instructions with human feedback.Advances in Neural Information Processing Systems, 35:27730–27744, 2022

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-09T16:18:40.522572Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T16:18:40.522572Z digest=sha256:86edea7a1181666e5e911ee8965665d5bc4f9fc024d941cc7c8417b569f7095c

Observation 648cc058-61a0-4053-8b17-978cb4b54180 · outbound

This paper cites Deep reinforcement learning from human preferences.Advances in neural information processing systems, 30, 2017.

On Almost Surely Safe Alignment of Large Language Models at Inference-Time Deep reinforcement learning from human preferences.Advances in neural information processing systems, 30, 2017

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-09T16:18:40.527252Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T16:18:40.527252Z digest=sha256:898964dbb2a683ba7664d9c80c0e0096be6a1fa869d3fc9e40ed4e5688e4964f

Observation dba0b472-7b72-4142-8613-083d21d44aca · outbound

This paper cites WebGPT: Browser-assisted question-answering with human feedback.

On Almost Surely Safe Alignment of Large Language Models at Inference-Time WebGPT: Browser-assisted question-answering with human feedback

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-09T16:18:40.531827Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T16:18:40.531827Z digest=sha256:f563305b262133d09b86968addf5ad2fd81560156a78d06dd1f76d75b7637f35

Observation 227a9f82-3d23-44eb-8d64-70512565a5d7 · outbound

This paper cites Controlled Decoding from Language Models.

On Almost Surely Safe Alignment of Large Language Models at Inference-Time Controlled Decoding from Language Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-09T16:18:40.536823Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T16:18:40.536823Z digest=sha256:3a397219a734e37629e6d47766970137fa54965404ed82c639767d1ab1444862

Observation 294b9ea5-5545-462b-97b3-54e10ec76f02 · outbound

This paper cites Discrete-time markov control processes with discounted unbounded costs: optimality criteria.Kybernetika, 28(3):191–212, 1992.

On Almost Surely Safe Alignment of Large Language Models at Inference-Time Discrete-time markov control processes with discounted unbounded costs: optimality criteria.Kybernetika, 28(3):191–212, 1992

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-09T16:18:40.542544Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T16:18:40.542544Z digest=sha256:a5a43bb66a095420dca408de69db55e4660bb0161a006942850b3d5b9ea3e495

Observation ee21386e-f652-4cd8-9501-9486c02f365e · outbound

This paper cites Sauté rl: Almost surely safe reinforcement learning using state augmentation.

On Almost Surely Safe Alignment of Large Language Models at Inference-Time Sauté rl: Almost surely safe reinforcement learning using state augmentation

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-09T16:18:40.547563Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T16:18:40.547563Z digest=sha256:783e18615c0e428ea1b000b6e4e1a48e1090c23460d5dd17aa35e050361469ce

Observation b7c85b2a-13af-4795-abf5-a1817f9f6c71 · outbound

This paper cites Fine-Tuning Language Models from Human Preferences.

On Almost Surely Safe Alignment of Large Language Models at Inference-Time Fine-Tuning Language Models from Human Preferences

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-09T16:18:40.552457Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T16:18:40.552457Z digest=sha256:291a5ab16d0bb6186284c6660fe72375c67e9bc55fa81bdfbc5a454c6943d374

Observation 3fd0f1f1-d309-4023-bd13-2b80c4327f83 · outbound

This paper cites Proximal Policy Optimization Algorithms.

On Almost Surely Safe Alignment of Large Language Models at Inference-Time Proximal Policy Optimization Algorithms

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-09T16:18:40.557404Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T16:18:40.557404Z digest=sha256:3823c7d34b536bdf44610f4a9ccf010450896c77bf667a859626900ca62a63a0

Observation dbf5f19f-ae5a-405f-8053-bb6ad02842ff · outbound

This paper cites Many of Your DPOs are Secretly One: Attempting Unification Through Mutual Information.

On Almost Surely Safe Alignment of Large Language Models at Inference-Time Many of Your DPOs are Secretly One: Attempting Unification Through Mutual Information

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-09T16:18:40.562457Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T16:18:40.562457Z digest=sha256:59e5bf4b16c33f8cfdfb738f9bfa93654c6aff2be89a5e2eb76de383eb36eda8

Observation 649dace3-01ac-4ac4-8cca-b3a0e6853cb3 · outbound

This paper cites Relative Preference Optimization: Enhancing LLM Alignment through Contrasting Responses across Identical and Diverse Prompts.

On Almost Surely Safe Alignment of Large Language Models at Inference-Time Relative Preference Optimization: Enhancing LLM Alignment through Contrasting Responses across Identical and Diverse Prompts

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-09T16:18:40.567531Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T16:18:40.567531Z digest=sha256:63329aa9eb28a6c5e31a6a333d249dad85b4369615607bba0ddcc47b6f5c3f16

Observation 90d13abe-0606-46a8-be83-baf1426b7be6 · outbound

This paper cites Direct Preference Optimization: Your Language Model is Secretly a Reward Model.

On Almost Surely Safe Alignment of Large Language Models at Inference-Time Direct Preference Optimization: Your Language Model is Secretly a Reward Model

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-09T16:18:40.572720Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T16:18:40.572720Z digest=sha256:b72afcf1c03719ef58e789e886eddbc098aaec7f5994f94840401fed39a0c3a8

Observation 7699aa7b-11f8-4f6e-8770-38fe3745ab72 · outbound

This paper cites A General Theoretical Paradigm to Understand Learning from Human Preferences.

On Almost Surely Safe Alignment of Large Language Models at Inference-Time A General Theoretical Paradigm to Understand Learning from Human Preferences

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-09T16:18:40.577727Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T16:18:40.577727Z digest=sha256:d90de6a348f5c648878f23e613fd888ae7a9d8cdb7e1b87d96bf262338caf599

Observation a6eccc06-70ba-46d1-a3e0-5158e1fbea3c · outbound

This paper cites SLiC-HF: Sequence Likelihood Calibration with Human Feedback.

On Almost Surely Safe Alignment of Large Language Models at Inference-Time SLiC-HF: Sequence Likelihood Calibration with Human Feedback

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-09T16:18:40.583638Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T16:18:40.583638Z digest=sha256:d260805cf1da4147caf2f0f2d3fffc6ffebd2d697b4c7a7292686209a4c52011

Observation 2cd7047a-8241-4f99-aed5-d87a2d392949 · outbound

This paper cites Generalized Preference Optimization: A Unified Approach to Offline Alignment.

On Almost Surely Safe Alignment of Large Language Models at Inference-Time Generalized Preference Optimization: A Unified Approach to Offline Alignment

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-09T16:18:40.588740Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T16:18:40.588740Z digest=sha256:b3199c24950de2a52e2b115eecc7421da55064ca8c5792774316323147af8ce7

Observation 6976cd11-331a-4878-8740-128ed3baadfe · outbound

This paper cites Preference ranking optimization for human alignment.

On Almost Surely Safe Alignment of Large Language Models at Inference-Time Preference ranking optimization for human alignment

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-09T16:18:40.593541Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T16:18:40.593541Z digest=sha256:eee18a924c87eb3509c9777a3ec79d5b2ae231aa0b2527f7cea17361643e192e

Observation 9213a5ed-efd0-4009-b833-fb1631a64bd1 · outbound

This paper cites KTO: Model Alignment as Prospect Theoretic Optimization.

On Almost Surely Safe Alignment of Large Language Models at Inference-Time KTO: Model Alignment as Prospect Theoretic Optimization

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-09T16:18:40.598165Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T16:18:40.598165Z digest=sha256:c62844187e476fa82aa84a7de10aac4090ce95df4840cc31ce0c7155ee4d46ed

Observation e5a525b3-1930-4ccc-8751-ff5f63ae4758 · outbound

This paper cites Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback.

On Almost Surely Safe Alignment of Large Language Models at Inference-Time Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-09T16:18:40.603344Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T16:18:40.603344Z digest=sha256:1fdad6c0ff23c8957316fc5eb55fb041d21ef61e06b831de47aeaf87d3dcce0c

Observation 6d24c066-c825-45cd-ab93-ea24dfd0afef · outbound

This paper cites Safe RLHF: Safe Reinforcement Learning from Human Feedback.

On Almost Surely Safe Alignment of Large Language Models at Inference-Time Safe RLHF: Safe Reinforcement Learning from Human Feedback

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-09T16:18:40.608158Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T16:18:40.608158Z digest=sha256:4bcee00af68c6d4f20b037286c8ee1f4efe8bbfd3fe7240a8738e933c3be3d30

Observation 96e435dd-c57c-48ea-9d6a-557397562cc5 · outbound

This paper cites Machine Unlearning in Large Language Models.

On Almost Surely Safe Alignment of Large Language Models at Inference-Time Machine Unlearning in Large Language Models

Reference 29

Resolution
verified exact
local_arxiv, observed 2026-08-09T16:18:42.415020Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-09T16:18:40.612953Z digest=sha256:15e83031219771d57b826defa61b7550326752c0033ab4af00c2a388ccad7266

Observation d5328590-8792-4f8a-98d4-bf55f78f9471 · outbound

This paper cites Eyes Closed, Safety On: Protecting Multimodal LLMs via Image-to-Text Transformation.

On Almost Surely Safe Alignment of Large Language Models at Inference-Time Eyes Closed, Safety On: Protecting Multimodal LLMs via Image-to-Text Transformation

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-09T16:18:40.617617Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T16:18:40.617617Z digest=sha256:ba38c22ed4c430a7b4131985571d15e6ecc5f39d9966e19386f96649dbccf657

Observation 572d086f-55e2-4102-8495-32dee9c94a4e · outbound

This paper cites Model Merging and Safety Alignment: One Bad Model Spoils the Bunch.

On Almost Surely Safe Alignment of Large Language Models at Inference-Time Model Merging and Safety Alignment: One Bad Model Spoils the Bunch

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-09T16:18:40.622233Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T16:18:40.622233Z digest=sha256:3ce8c760339a49b46a95f16e16cef85344d9197600881a0cfba04d0ab9d7a8ce

Observation b31caba9-2112-4e1f-85d9-8ba70e4cf7e2 · outbound

This paper cites Trustagent: Towards safe and trustworthy llm-based agents through agent constitution.

On Almost Surely Safe Alignment of Large Language Models at Inference-Time Trustagent: Towards safe and trustworthy llm-based agents through agent constitution

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-09T16:18:40.626509Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T16:18:40.626509Z digest=sha256:f7df51c9fd316cc6bd2c36584c1a4f13356a3fc5ffa836b15e06c4e3f029f8da

Observation 396404a1-a9e6-456f-957b-347827e1b2d8 · outbound

This paper cites Controllable Safety Alignment: Inference-Time Adaptation to Diverse Safety Requirements.

On Almost Surely Safe Alignment of Large Language Models at Inference-Time Controllable Safety Alignment: Inference-Time Adaptation to Diverse Safety Requirements

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-09T16:18:40.631160Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T16:18:40.631160Z digest=sha256:677a8e37d9fe748a3f64fddde021154c7cf76e245893edc6e3ff0bd1aba88222

Observation 4bcea3a2-2472-4028-ba55-23a4c040a8a4 · outbound

This paper cites COLD-Attack: Jailbreaking LLMs with Stealthiness and Controllability.

On Almost Surely Safe Alignment of Large Language Models at Inference-Time COLD-Attack: Jailbreaking LLMs with Stealthiness and Controllability

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-09T16:18:40.636116Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T16:18:40.636116Z digest=sha256:62296675febfa744a6001d49ae63bb2aa406a307db5ad6def2a6701ab4d8f0c0

Observation bc805eee-7db5-4c63-9bc9-dcad2a38642b · outbound

This paper cites SafeDecoding: Defending against Jailbreak Attacks via Safety-Aware Decoding.

On Almost Surely Safe Alignment of Large Language Models at Inference-Time SafeDecoding: Defending against Jailbreak Attacks via Safety-Aware Decoding

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-09T16:18:40.641145Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T16:18:40.641145Z digest=sha256:a1b5b8c105ceec947623d2037e73ac9203fb9cfabfc39fed224350d3444261b5

Observation 37ee5772-0a77-4982-9ca0-9adb1b4979c5 · outbound

This paper cites Assessing the Brittleness of Safety Alignment via Pruning and Low-Rank Modifications.

On Almost Surely Safe Alignment of Large Language Models at Inference-Time Assessing the Brittleness of Safety Alignment via Pruning and Low-Rank Modifications

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-09T16:18:40.645718Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T16:18:40.645718Z digest=sha256:7462808775b9df94204a3f7fef220eef6a6cfb4407f1d35cd7641a482531334d

Observation 43d54e6f-706a-475d-833c-5df61b19afa8 · outbound

This paper cites SaLoRA: Safety-Alignment Preserved Low-Rank Adaptation.

On Almost Surely Safe Alignment of Large Language Models at Inference-Time SaLoRA: Safety-Alignment Preserved Low-Rank Adaptation

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-09T16:18:40.650485Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T16:18:40.650485Z digest=sha256:e9957f0501d71fac9c5cc2311b1264ad69ff36add560dcdce470dc1897e728a7

Observation 421ee288-8eca-40b0-bcba-50bee8ed51d5 · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

On Almost Surely Safe Alignment of Large Language Models at Inference-Time Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-09T16:18:40.655310Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T16:18:40.655310Z digest=sha256:de835e14c6c715e1b31bd9b135a659c5d32f395daa200bbe23b85bf12cec86d5

Observation 1da9e745-d102-4f7f-a36f-f3ae579b8f93 · outbound

This paper cites Fast Best-of-N Decoding via Speculative Rejection.

On Almost Surely Safe Alignment of Large Language Models at Inference-Time Fast Best-of-N Decoding via Speculative Rejection

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-09T16:18:40.660327Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T16:18:40.660327Z digest=sha256:74c1b9da23a59597b7745e6ff84347a0e7b0b5b1cbac3adfc4932409b265dfa1

Observation e5f58d55-9049-4992-b048-707d9eb6478e · outbound

This paper cites Fudge: Controlled text generation with future discriminators.

On Almost Surely Safe Alignment of Large Language Models at Inference-Time Fudge: Controlled text generation with future discriminators

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-09T16:18:40.665059Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T16:18:40.665059Z digest=sha256:e51516eb263adba3fed15bd3bad0e47330836679fdbba365337ec4b5977d097f

Observation 10373a1d-8d46-4e0d-af28-550656ed2934 · outbound

This paper cites Cold decoding: Energy- based constrained text generation with langevin dynamics.Advances in Neural Information Processing Systems, 35:9538–9551, 2022.

On Almost Surely Safe Alignment of Large Language Models at Inference-Time Cold decoding: Energy- based constrained text generation with langevin dynamics.Advances in Neural Information Processing Systems, 35:9538–9551, 2022

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-09T16:18:40.669670Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T16:18:40.669670Z digest=sha256:e5c0012bfa1bfaef8b86fd43db40e5c6aefb923030167ad1976bef5257d4a9c7

Observation 533b95a6-6aa5-4fff-a450-6c09a6dc7d10 · outbound

This paper cites Aligning Large Language Models with Representation Editing: A Control Perspective.

On Almost Surely Safe Alignment of Large Language Models at Inference-Time Aligning Large Language Models with Representation Editing: A Control Perspective

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-09T16:18:40.674151Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T16:18:40.674151Z digest=sha256:48d1cb312fa352b0dfaeb930712afaca195d6a3914517b06549b9cc3a9257ed2

Observation 124f1ebe-35b4-4a99-8e30-73cfad6e3a45 · outbound

This paper cites ARGS: Alignment as Reward-Guided Search.

On Almost Surely Safe Alignment of Large Language Models at Inference-Time ARGS: Alignment as Reward-Guided Search

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-09T16:18:40.679002Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T16:18:40.679002Z digest=sha256:93e9ea400645b2fd95091af013b81ce72bf37def1bf0e88c7a23eb24c8503566

Observation 208bf629-eee1-486a-83d3-f438ff773a75 · outbound

This paper cites Decoding-Time Language Model Alignment with Multiple Objectives.

On Almost Surely Safe Alignment of Large Language Models at Inference-Time Decoding-Time Language Model Alignment with Multiple Objectives

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-09T16:18:40.683972Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T16:18:40.683972Z digest=sha256:d28781eea0c37740b08ab9ccd5e0d8aacb93de9c0cebb28dbd8bbea60b13d14a

Observation 766529f0-d5ea-44f3-9022-4364b48cfb8e · outbound

This paper cites Deal: Decoding-time alignment for large language models.arXiv preprint arXiv:2402.06147, 2024.

On Almost Surely Safe Alignment of Large Language Models at Inference-Time Deal: Decoding-time alignment for large language models.arXiv preprint arXiv:2402.06147, 2024

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-09T16:18:40.688825Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T16:18:40.688825Z digest=sha256:8fee193332777c66136c867cd3c403783cfabd9d95a8f63e18c2d702da486886

Observation 8714e22e-567f-4363-9f27-0c577114942b · outbound

This paper cites Value Augmented Sampling for Language Model Alignment and Personalization.

On Almost Surely Safe Alignment of Large Language Models at Inference-Time Value Augmented Sampling for Language Model Alignment and Personalization

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-09T16:18:40.693492Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T16:18:40.693492Z digest=sha256:c05e56a891c814b80aa6e3b38e2bd298e3898129795ce26deeb13d1956ebf1b9

Observation 123dee81-6968-497b-8966-f3b1c026d33a · outbound

This paper cites ROSE Doesn't Do That: Boosting the Safety of Instruction-Tuned Large Language Models with Reverse Prompt Contrastive Decoding.

On Almost Surely Safe Alignment of Large Language Models at Inference-Time ROSE Doesn't Do That: Boosting the Safety of Instruction-Tuned Large Language Models with Reverse Prompt Contrastive Decoding

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-09T16:18:40.699323Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T16:18:40.699323Z digest=sha256:e2900fe6cc9477d2c443b086d7464aa2175876b06c7f55fa22efdbe5875cc56a

Observation 5830fdfc-6502-443e-a20d-d3c788e9e189 · outbound

This paper cites Adversarial contrastive decoding: Boosting safety alignment of large language models via opposite prompt optimization.arXiv preprint arXiv:2406.16743, 2024.

On Almost Surely Safe Alignment of Large Language Models at Inference-Time Adversarial contrastive decoding: Boosting safety alignment of large language models via opposite prompt optimization.arXiv preprint arXiv:2406.16743, 2024

Reference 48

Resolution
verified exact
raw_fallback, observed 2026-08-09T16:18:42.059810Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-09T16:18:40.704343Z digest=sha256:e724c1554ac4cee93a1b4535cb3991f30a7d8fe4e07bcab3ec667d9d72e88cab

Observation e17d599a-d879-4a45-9d9f-34a81585f77f · outbound

This paper cites Parameter-Efficient Detoxification with Contrastive Decoding.

On Almost Surely Safe Alignment of Large Language Models at Inference-Time Parameter-Efficient Detoxification with Contrastive Decoding

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-09T16:18:40.709137Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T16:18:40.709137Z digest=sha256:be42fb12c2003172ad9ac3c2110d4671c6f2f1560619c7a57d7afc02a0a5ef44

Observation 4c174fe1-5db0-4db6-bd00-cded87ecad8d · outbound

This paper cites Root Defence Strategies: Ensuring Safety of LLM at the Decoding Level.

On Almost Surely Safe Alignment of Large Language Models at Inference-Time Root Defence Strategies: Ensuring Safety of LLM at the Decoding Level

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-09T16:18:40.714261Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T16:18:40.714261Z digest=sha256:f6747694539c47eea065cc4c7d9f89decac12b0960af0b07809361340059b6ec

Observation 92d87873-ae93-4ffb-9f2c-8da28b49f108 · outbound

This paper cites Attacks, Defenses and Evaluations for LLM Conversation Safety: A Survey.

On Almost Surely Safe Alignment of Large Language Models at Inference-Time Attacks, Defenses and Evaluations for LLM Conversation Safety: A Survey

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-09T16:18:40.719150Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T16:18:40.719150Z digest=sha256:21a66815f728f78368800dc2f32d987dbbe6e9e598c147eaed99b8c45be8ef90

Observation 9faece8a-cb64-4cee-8fae-f19b7dd27043 · outbound

This paper cites Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations.

On Almost Surely Safe Alignment of Large Language Models at Inference-Time Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-09T16:18:40.723901Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T16:18:40.723901Z digest=sha256:1e7e1a5efa378bf7fefcd643b96027d13684433acf1151256c3c0ceb59befec8

Observation f656bf9e-92b7-42a5-bae9-d40de150813e · outbound

This paper cites Probing the Safety Response Boundary of Large Language Models via Unsafe Decoding Path Generation.

On Almost Surely Safe Alignment of Large Language Models at Inference-Time Probing the Safety Response Boundary of Large Language Models via Unsafe Decoding Path Generation

Reference 53

Resolution
verified exact
local_arxiv, observed 2026-08-09T16:18:41.914435Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-09T16:18:40.728371Z digest=sha256:cb0a3e382332d83b4f38492544b91f0453ec456c9722d948d634e83f2fa3e1ac

Observation 1f8db701-1072-47cc-8803-4c7c0730eb18 · outbound

This paper cites Mixture of attentions for speculative decoding, 2024.

On Almost Surely Safe Alignment of Large Language Models at Inference-Time Mixture of attentions for speculative decoding, 2024

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-09T16:18:40.732968Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T16:18:40.732968Z digest=sha256:c6a71efbc166c868743a904816ce9e7f489e3b393cbd6eb8d6102052800b3b00

Observation 651c2a9e-9828-4a4a-bdb5-73a55ba1f21c · outbound

This paper cites DPO Meets PPO: Reinforced Token Optimization for RLHF.

On Almost Surely Safe Alignment of Large Language Models at Inference-Time DPO Meets PPO: Reinforced Token Optimization for RLHF

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-09T16:18:40.737239Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T16:18:40.737239Z digest=sha256:5ade489a0eed0e1579ecbf47102e75bfff79132eac9c9493d749ed24ac2f06d7

Observation ed3494c6-c7ef-4308-aeb9-2c7b04a6a608 · outbound

This paper cites Large language models for robotics: A survey.arXiv preprint arXiv:2311.07226, 2023.

On Almost Surely Safe Alignment of Large Language Models at Inference-Time Large language models for robotics: A survey.arXiv preprint arXiv:2311.07226, 2023

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-09T16:18:40.741644Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T16:18:40.741644Z digest=sha256:455ae5741a2b1e2feda36bc154cfd1786bed006d721145c29e34b80e727181ef

Observation 7145db43-1cca-4f4d-ab05-92a39914a809 · outbound

This paper cites Leandojo: Theorem proving with retrieval-augmented language models.Advances in Neural Information Processing Systems, 36, 2024.

On Almost Surely Safe Alignment of Large Language Models at Inference-Time Leandojo: Theorem proving with retrieval-augmented language models.Advances in Neural Information Processing Systems, 36, 2024

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-09T16:18:40.746544Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T16:18:40.746544Z digest=sha256:ae63def433e8416d605640d4d72b317b0a46612b37254ea5cbb01a73669a4cc9

Observation b1b9a395-8e6f-4b0b-a60b-482ac08faa59 · outbound

This paper cites Solving olympiad geometry without human demonstrations.Nature, 625(7995):476–482, 2024.

On Almost Surely Safe Alignment of Large Language Models at Inference-Time Solving olympiad geometry without human demonstrations.Nature, 625(7995):476–482, 2024

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-09T16:18:40.751431Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T16:18:40.751431Z digest=sha256:683d238f4a014b9c30795c37648a2ba08cc82d59ae88ff5f48f878bf0aa24ed1

Observation da1743d3-e76f-4297-bd1c-4beeb9758733 · outbound

This paper cites Learn from Failure: Fine-Tuning LLMs with Trial-and-Error Data for Intuitionistic Propositional Logic Proving.

On Almost Surely Safe Alignment of Large Language Models at Inference-Time Learn from Failure: Fine-Tuning LLMs with Trial-and-Error Data for Intuitionistic Propositional Logic Proving

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-09T16:18:40.756147Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T16:18:40.756147Z digest=sha256:3a7e2531d981e0480079c62813cc1182dde7dfbe622d57dbe7092ad5308b3ad3

Observation 9e5ba516-0a06-4f77-9154-6f82a0b8c9c8 · outbound

This paper cites Learning to Learn Faster from Human Feedback with Language Model Predictive Control.

On Almost Surely Safe Alignment of Large Language Models at Inference-Time Learning to Learn Faster from Human Feedback with Language Model Predictive Control

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-09T16:18:40.761602Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T16:18:40.761602Z digest=sha256:b79a9d41ca498775d1454b6abeb1efe3ab9fee72f9fd95629ffd5fdcb85d10ea

Observation 5705fb0c-2d13-4b52-80af-c767cd16424e · outbound

This paper cites Math-shepherd: Verify and reinforce llms step-by-step without human annotations.

On Almost Surely Safe Alignment of Large Language Models at Inference-Time Math-shepherd: Verify and reinforce llms step-by-step without human annotations

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-09T16:18:40.766715Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T16:18:40.766715Z digest=sha256:18eb5f5ba67aa7088b8b11e46048bc3f5a753f0126932bfe714ba3ecdfbfd56b

Observation 42aaf597-5211-412e-b112-53e20c5aebc7 · outbound

This paper cites ReST-MCTS*: LLM Self-Training via Process Reward Guided Tree Search.

On Almost Surely Safe Alignment of Large Language Models at Inference-Time ReST-MCTS*: LLM Self-Training via Process Reward Guided Tree Search

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-09T16:18:40.771963Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T16:18:40.771963Z digest=sha256:a672aad349df7433d496d4e92ae2227e6bacfd913ac53c9537ecc5e173fe6f24

Observation ac704221-d03e-4dad-b832-cde745c3475e · outbound

This paper cites Enhancing Reasoning through Process Supervision with Monte Carlo Tree Search.

On Almost Surely Safe Alignment of Large Language Models at Inference-Time Enhancing Reasoning through Process Supervision with Monte Carlo Tree Search

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-09T16:18:40.777676Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T16:18:40.777676Z digest=sha256:b49e0f1964b08a8a1fc7bb9c62323848dd29c0feeffbfe1c968660f0c17a5855

Observation bef3432f-ac89-4a0e-9fd2-fcfa2702a5d4 · outbound

This paper cites Simulation-guided beam search for neural combinatorial optimization.

On Almost Surely Safe Alignment of Large Language Models at Inference-Time Simulation-guided beam search for neural combinatorial optimization

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-09T16:18:40.782889Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T16:18:40.782889Z digest=sha256:40858459899dd756e199dffef36bd8797ddcf63362e25e48904e97feb3d410ca

Observation e3d415a0-4764-41cb-812e-69dcc65e9d12 · outbound

This paper cites Constrained policy optimization.

On Almost Surely Safe Alignment of Large Language Models at Inference-Time Constrained policy optimization

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-09T16:18:40.788403Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T16:18:40.788403Z digest=sha256:6e49b8599bb8583ef55e1614ff12011ed0be77be3fecb07da6cd21f75b9c35bd

Observation de7ee1c1-9ed3-4adb-86e6-a016d3d0146e · outbound

This paper cites A Survey on LLM Test-Time Compute via Search: Tasks, LLM Profiling, Search Algorithms, and Relevant Frameworks.

On Almost Surely Safe Alignment of Large Language Models at Inference-Time A Survey on LLM Test-Time Compute via Search: Tasks, LLM Profiling, Search Algorithms, and Relevant Frameworks

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-09T16:18:40.793196Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T16:18:40.793196Z digest=sha256:9a2c0cf4729faa09053db9850b19868aae98598627d761636abd6eca7accd442

Observation e5619f61-9495-4266-a530-23b98cdca7b0 · outbound

This paper cites Alpaca: A strong, replicable instruction- following model.Stanford Center for Research on Foundation Models.

On Almost Surely Safe Alignment of Large Language Models at Inference-Time Alpaca: A strong, replicable instruction- following model.Stanford Center for Research on Foundation Models

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-09T16:18:40.798466Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T16:18:40.798466Z digest=sha256:a80a83f9bcb05e84653295df361009e7828b71df3a43af4198d9cd4872f62c5a

Observation b7de5c36-f360-4a9f-b87f-ef62f7b04c7d · outbound

This paper cites Gonzalez, Ion Stoica, and Eric P.

On Almost Surely Safe Alignment of Large Language Models at Inference-Time Gonzalez, Ion Stoica, and Eric P

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-09T16:18:40.803395Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T16:18:40.803395Z digest=sha256:4f2fc1d2aff7ae2bf3cbaa7cc5e2aa51fe8c705ee97e88f40ece194c87a52aa9

Observation 7d07096d-3017-4fed-b163-1c123b3a5bbc · outbound

This paper cites The Llama 3 Herd of Models.

On Almost Surely Safe Alignment of Large Language Models at Inference-Time The Llama 3 Herd of Models

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-09T16:18:40.808370Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T16:18:40.808370Z digest=sha256:7af8f9d5843c3178fcad862a66e31274ae86571b4ea4bc9f18bb4412d7262788

Observation e12b652f-af52-4be7-b387-036e5cfe24d3 · outbound

This paper cites PKU-SafeRLHF: Towards Multi-Level Safety Alignment for LLMs with Human Preference.

On Almost Surely Safe Alignment of Large Language Models at Inference-Time PKU-SafeRLHF: Towards Multi-Level Safety Alignment for LLMs with Human Preference

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-09T16:18:40.813673Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T16:18:40.813673Z digest=sha256:bde667773c2112a16f9fe90997222f2531545faebd97f8d53b336c1e3c5a70f0

Observation 5d4c7c56-409a-4a5e-9726-ca7b20976493 · outbound

This paper cites Quantile Regression for Distributional Reward Models in RLHF.

On Almost Surely Safe Alignment of Large Language Models at Inference-Time Quantile Regression for Distributional Reward Models in RLHF

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-09T16:18:40.819124Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T16:18:40.819124Z digest=sha256:ea1fdb95a392775dbd78eb1744c14a4457e02ce7f320fae742f78a17ba26c929

Observation 0869e124-9f78-413a-82c8-6378656a10ec · outbound

This paper cites CRC Press, 1999.

On Almost Surely Safe Alignment of Large Language Models at Inference-Time CRC Press, 1999

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-09T16:18:40.824171Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T16:18:40.824171Z digest=sha256:9327cedca925f7729442af9cf073e33c1060036a59a7a009f591af82e3b67117

Observation 5b182e98-1916-42fa-bce7-c4ad533480d6 · outbound

This paper cites Safe exploration in finite markov decision processes with gaussian processes.

On Almost Surely Safe Alignment of Large Language Models at Inference-Time Safe exploration in finite markov decision processes with gaussian processes

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-09T16:18:40.829059Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T16:18:40.829059Z digest=sha256:66ef98f769a80978c4352a9319b76dc6a6101387df340e94daf5cf1bd2591255

Observation 1bf01d93-0d87-42ce-a967-2632b8e3c341 · outbound

This paper cites Learning-based model predictive control for safe exploration.

On Almost Surely Safe Alignment of Large Language Models at Inference-Time Learning-based model predictive control for safe exploration

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-09T16:18:40.834159Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T16:18:40.834159Z digest=sha256:5413a9ed206bfcb42ac9eea1a6655ad87982c9cc5e4cdadf332bccb1d498a594

Observation 6c6c3dfd-6153-4bda-ab02-ec7d776a2e50 · outbound

This paper cites Safe exploration in continuous action spaces.

On Almost Surely Safe Alignment of Large Language Models at Inference-Time Safe exploration in continuous action spaces

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-09T16:18:40.839073Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T16:18:40.839073Z digest=sha256:29d42d00044a6187d86f64aac4115a20e1d2773d5f0c9a2fbf4a1df3dbb62abe

Observation 5634e76d-a329-41bf-bd30-0d3f27d24f45 · outbound

This paper cites Safe exploration and optimization of constrained mdps using gaussian processes.

On Almost Surely Safe Alignment of Large Language Models at Inference-Time Safe exploration and optimization of constrained mdps using gaussian processes

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-09T16:18:40.844031Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T16:18:40.844031Z digest=sha256:d48c5d747c43c1b12fd7bdff2d635c85d7b0c1387f4c9b687442dd9105693c63

Observation c1d9e1c2-9de4-4dca-b490-7b4ede335b4e · outbound

This paper cites Conservative safety critics for exploration.

On Almost Surely Safe Alignment of Large Language Models at Inference-Time Conservative safety critics for exploration

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-09T16:18:40.849005Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T16:18:40.849005Z digest=sha256:58c928d2f9c6ef31a8f91842e3bd207c45bb67bd9a8cdd14a40bef8bd170bcd8

Observation 0d2d19ca-97bd-4922-9920-8a2271f945fb · outbound

This paper cites Lyapunov-based safe policy optimization for continuous control.

On Almost Surely Safe Alignment of Large Language Models at Inference-Time Lyapunov-based safe policy optimization for continuous control

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-09T16:18:40.853910Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T16:18:40.853910Z digest=sha256:fb335863a15b8baaa8827d03800a7cc5e3aa7256b563eae783ccc032a926c47d

Observation 64b53dcf-dbfb-4ebb-85cc-673393c5aa59 · outbound

This paper cites Lyapunov-based Safe Policy Optimization for Continuous Control.

On Almost Surely Safe Alignment of Large Language Models at Inference-Time Lyapunov-based Safe Policy Optimization for Continuous Control

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-09T16:18:40.859342Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T16:18:40.859342Z digest=sha256:bf8dbf1a5ce2b44f215cc7742de5b58f089c0c4e6b148c66be204bca22ba0813

Observation 95d7b4cf-1ca3-4cdc-ab9a-8240afb01643 · outbound

This paper cites Safe model- based reinforcement learning with stability guarantees.

On Almost Surely Safe Alignment of Large Language Models at Inference-Time Safe model- based reinforcement learning with stability guarantees

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-09T16:18:40.864852Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T16:18:40.864852Z digest=sha256:1134a565ae7dbca12b8640b8314f8f5020db4afe60b4fa71dc4d40ec821780bb

Observation ac53c036-895a-4b39-9feb-c7b3807c3c44 · outbound

This paper cites Barrier-certified adaptive reinforcement learning with applications to brushbot navigation.

On Almost Surely Safe Alignment of Large Language Models at Inference-Time Barrier-certified adaptive reinforcement learning with applications to brushbot navigation

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-09T16:18:40.869206Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T16:18:40.869206Z digest=sha256:fd4f67f5cc36f46968bb7a45d78cb79a06773dc1ce648718034d60046a863556

Observation 7d4be5b3-756c-43d3-8f69-6bf505e44aab · outbound

This paper cites End-to-end safe reinforcement learning through barrier functions for safety-critical continuous control tasks.

On Almost Surely Safe Alignment of Large Language Models at Inference-Time End-to-end safe reinforcement learning through barrier functions for safety-critical continuous control tasks

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-09T16:18:40.873885Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T16:18:40.873885Z digest=sha256:64423f16ce379ce2c8886bda879d03245f9b12f1caccca0dd9d32dd2a7cab616

Observation 4ae5cb98-1c8d-4ab4-b423-a7668f7417e0 · outbound

This paper cites Reachability-based safe learning with gaussian processes.

On Almost Surely Safe Alignment of Large Language Models at Inference-Time Reachability-based safe learning with gaussian processes

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-09T16:18:40.878073Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T16:18:40.878073Z digest=sha256:408a2f929419781ba3f8ac12d1d65c7914ceb7a35c6ff51afd6fda588593547d

Observation ef3f6ca5-4277-4b1d-8799-272cf280e262 · outbound

This paper cites Safeguarding resource- constrained cyber-physical systems with adaptive control.

On Almost Surely Safe Alignment of Large Language Models at Inference-Time Safeguarding resource- constrained cyber-physical systems with adaptive control

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-09T16:18:40.882265Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T16:18:40.882265Z digest=sha256:409bc4693377af2dbfdec92029355fd6935ef4e37d546189590a7048101099b6

Observation 6696e5ab-2224-4843-b2df-141817be418d · outbound

This paper cites Bridging model-based safety and model-free reinforcement learning through system identification and safety-critical control.

On Almost Surely Safe Alignment of Large Language Models at Inference-Time Bridging model-based safety and model-free reinforcement learning through system identification and safety-critical control

Reference 85

Resolution
unresolved
no resolver link, observed 2026-08-09T16:18:40.886651Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T16:18:40.886651Z digest=sha256:7faa77c117f9846087377eef09b03e64a409bbbe701b9842331aac9d0ecfbeee

Observation 8ab2b39d-5b6a-4d8f-9c07-73f538ae03a1 · outbound

This paper cites Benchmarking safe exploration in deep reinforcement learning.

On Almost Surely Safe Alignment of Large Language Models at Inference-Time Benchmarking safe exploration in deep reinforcement learning

Reference 86

Resolution
unresolved
no resolver link, observed 2026-08-09T16:18:40.890979Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T16:18:40.890979Z digest=sha256:77b14e9e326d198c852ff7c2c9fd3c11f3a43cddbb9d76b8adc2dc397782b42b

Observation 8ef6e941-b931-4824-9cb7-770e74557421 · outbound

This paper cites Responsive safety in reinforcement learning by monitoring risk and adapting policies.

On Almost Surely Safe Alignment of Large Language Models at Inference-Time Responsive safety in reinforcement learning by monitoring risk and adapting policies

Reference 87

Resolution
unresolved
no resolver link, observed 2026-08-09T16:18:40.895451Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T16:18:40.895451Z digest=sha256:d46520498fbe1b09c400e788fd9b23ff202d387cdcfe1477fe7c51bee789078c

Observation 50a8bd5b-4ba4-4050-bcad-7497e9c4a49a · outbound

This paper cites Relative value learning for constrained reinforce- ment learning.

On Almost Surely Safe Alignment of Large Language Models at Inference-Time Relative value learning for constrained reinforce- ment learning

Reference 88

Resolution
unresolved
no resolver link, observed 2026-08-09T16:18:40.899713Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T16:18:40.899713Z digest=sha256:4717bcb7527a80a19ae0a67a5d36448df05f97393bdabb48604fa2ea19fc682e

Observation 1aaaa581-e9b4-4bb2-a00b-f002d65c6940 · outbound

This paper cites Natural policy gradient for safe reinforcement learning with c-mdps.

On Almost Surely Safe Alignment of Large Language Models at Inference-Time Natural policy gradient for safe reinforcement learning with c-mdps

Reference 89

Resolution
unresolved
no resolver link, observed 2026-08-09T16:18:40.904474Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T16:18:40.904474Z digest=sha256:d5739e64306ad29e63f132f16c73a3e5149abe35922f8f61e2059201658de581

Observation 5f4be3f3-1c0b-4456-b84c-0429c3c7a294 · outbound

This paper cites Group Robust Preference Optimization in Reward-free RLHF.

On Almost Surely Safe Alignment of Large Language Models at Inference-Time Group Robust Preference Optimization in Reward-free RLHF

Reference 90

Resolution
unresolved
no resolver link, observed 2026-08-09T16:18:40.909170Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T16:18:40.909170Z digest=sha256:e6a80685a4b7d4f634675a9114a5c01b87b72b1f952c3da1d83dc5d76bb8ea54

Observation ced521c4-6a4b-4a5f-a0ec-90a99883d58a · outbound

This paper cites Mission Impossible: A Statistical Perspective on Jailbreaking LLMs.

On Almost Surely Safe Alignment of Large Language Models at Inference-Time Mission Impossible: A Statistical Perspective on Jailbreaking LLMs

Reference 91

Resolution
unresolved
no resolver link, observed 2026-08-09T16:18:40.914123Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T16:18:40.914123Z digest=sha256:3afc5c9e4e26df6537c1cc412cfaa92f6d9f78ac14ca4c82090167015e247aa8

Observation f4af410e-f2c1-4553-8007-fda5028b0cf7 · outbound

This paper cites Improving llm safety alignment with dual-objective optimization.arXiv preprint arXiv:2503.03710, 2025.

On Almost Surely Safe Alignment of Large Language Models at Inference-Time Improving llm safety alignment with dual-objective optimization.arXiv preprint arXiv:2503.03710, 2025

Reference 92

Resolution
unresolved
no resolver link, observed 2026-08-09T16:18:40.919238Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T16:18:40.919238Z digest=sha256:3003c9f107005345d44e089f565e45acecf0ab3991745d7f136e066e3e7c4f70

Observation cba5d427-4d84-4a84-a1c6-7d81a20c60d7 · outbound

This paper cites Safety Arithmetic: A Framework for Test-time Safety Alignment of Language Models by Steering Parameters and Activations.

On Almost Surely Safe Alignment of Large Language Models at Inference-Time Safety Arithmetic: A Framework for Test-time Safety Alignment of Language Models by Steering Parameters and Activations

Reference 94

Resolution
unresolved
no resolver link, observed 2026-08-09T16:18:40.934070Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T16:18:40.934070Z digest=sha256:b41b41dddc68ee93b0751a7bc13232cc57a25101aa646226c38eac506fb76f85

Observation ca5bb15d-04cd-456c-8255-defec97a947e · outbound

This paper cites On prompt-driven safeguarding for large language models.

On Almost Surely Safe Alignment of Large Language Models at Inference-Time On prompt-driven safeguarding for large language models

Reference 95

Resolution
unresolved
no resolver link, observed 2026-08-09T16:18:40.939000Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T16:18:40.939000Z digest=sha256:41686dc875ad0c62e19813369e3a3568d061118bd3b2f4e573c5ba4b7f7c1bc5

Observation 834e3454-ed9c-4937-b478-b541885fde2f · outbound

This paper cites Dynamic Guided and Domain Applicable Safeguards for Enhanced Security in Large Language Models.

On Almost Surely Safe Alignment of Large Language Models at Inference-Time Dynamic Guided and Domain Applicable Safeguards for Enhanced Security in Large Language Models

Reference 96

Resolution
verified exact
local_arxiv, observed 2026-08-09T16:18:41.475394Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-09T16:18:40.944028Z digest=sha256:25b1b8191aeb19e166925d53ef307830ee4180f2c1e605059c23744176080254

Observation 9bf19622-0f6e-4d67-9f43-8ab2e04fed6c · outbound

This paper cites Safety Alignment Should Be Made More Than Just a Few Tokens Deep.

On Almost Surely Safe Alignment of Large Language Models at Inference-Time Safety Alignment Should Be Made More Than Just a Few Tokens Deep

Reference 97

Resolution
unresolved
no resolver link, observed 2026-08-09T16:18:40.949223Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T16:18:40.949223Z digest=sha256:25dd2b5629e12d511b065f9e3f000bc7e3dd1f23d044bb28f3f6744e4eaa27f8

Observation a7ec959a-c7c2-4bb0-bdc7-2514815d82f6 · outbound

This paper cites Shaping the Safety Boundaries: Understanding and Defending Against Jailbreaks in Large Language Models.

On Almost Surely Safe Alignment of Large Language Models at Inference-Time Shaping the Safety Boundaries: Understanding and Defending Against Jailbreaks in Large Language Models

Reference 98

Resolution
unresolved
no resolver link, observed 2026-08-09T16:18:40.954512Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T16:18:40.954512Z digest=sha256:117d3a8809c1a415fd29baae6b4272a0bfac7bd3aaaf2afca84f0fbae0386508

Observation bdbb141f-9b63-42cf-95ec-a002ad5818ce · outbound

This paper cites DYNASHIELD: A Black-Box Moving Target Defense for LLMs via Dynamic Decoding Customization.

On Almost Surely Safe Alignment of Large Language Models at Inference-Time DYNASHIELD: A Black-Box Moving Target Defense for LLMs via Dynamic Decoding Customization

Reference 99

Resolution
unresolved
no resolver link, observed 2026-08-09T16:18:40.959416Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T16:18:40.959416Z digest=sha256:ccbb631334d3b040a7d68dd714679c8ef94ea8227e9ef7f6a22408c50c8b9bae

Observation e6d9fe43-9f48-4960-bd25-9b2c985a3985 · outbound

This paper cites Chain-of-detection enables robust and efficient jailbreak defense.

On Almost Surely Safe Alignment of Large Language Models at Inference-Time Chain-of-detection enables robust and efficient jailbreak defense

Reference 100

Resolution
unresolved
no resolver link, observed 2026-08-09T16:18:40.964386Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T16:18:40.964386Z digest=sha256:5dc9a5353a7e5e9f90d76d6aab6794b83e73c1d3931c57c3cd771c8cb455b111

Observation 7f5930cc-e48e-4a7f-99cc-629663c6796f · outbound

This paper cites Prefix Guidance: A Steering Wheel for Large Language Models to Defend Against Jailbreak Attacks.

On Almost Surely Safe Alignment of Large Language Models at Inference-Time Prefix Guidance: A Steering Wheel for Large Language Models to Defend Against Jailbreak Attacks

Reference 101

Resolution
verified exact
local_arxiv, observed 2026-08-09T16:18:41.400163Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-09T16:18:40.969178Z digest=sha256:78c28c08cc68d20739e380b9714991aee881dadad547918bf1511ffa2d35fe11

Pith citing papers

Observation 101ec2a0-0806-42e6-897c-69a238407f5e · inbound

Generative AI Act II: Test Time Scaling Drives Cognition Engineering cites this paper.

Generative AI Act II: Test Time Scaling Drives Cognition Engineering On Almost Surely Safe Alignment of Large Language Models at Inference-Time

Reference 136

Resolution
unresolved
no resolver link, observed 2026-08-16T12:02:43.148072Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T12:02:43.148072Z digest=sha256:fea621e64f9dbaf6b25fd27b73c78d1d4825013b4eee162143da8217f583ca39

Observation f47fd653-d640-4acc-b755-e93a1f270232 · inbound

Personalized Constitutionally-Aligned Agentic Superego: Secure AI Behavior Aligned to Diverse Human Values cites this paper.

Personalized Constitutionally-Aligned Agentic Superego: Secure AI Behavior Aligned to Diverse Human Values On Almost Surely Safe Alignment of Large Language Models at Inference-Time

Reference 31

Resolution
verified exact
local_arxiv, observed 2026-08-07T05:43:07.630186Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T05:43:07.249475Z digest=sha256:d37b3a95fef4c47f253aed0f6cb5f898fdcc9105152aca174f19ceed9f633bff