Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-10T22:35:27.013254Z
Paper Citation Record · LEDGER
As of 11 August 2026, this Paper Citation Record lists 28 of 28 outbound references and 3 inbound Pith citation observations for arXiv:2501.02018.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-10T22:35:27.013254Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-09T16:18:41.044579Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-08-05T12:54:07.656435Z
28 of 28 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation b11812df-a75e-499f-a9e1-35055655b0f2 · outbound
Safeguarding Large Language Models in Real-time with Tunable Safety-Performance Trade-offs Foundational Challenges in Assuring Alignment and Safety of Large Language Models
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 559dfc18-6046-4501-81c5-960609ef393e · outbound
Safeguarding Large Language Models in Real-time with Tunable Safety-Performance Trade-offs Nudging: Inference-time Alignment of LLMs via Guided Decoding
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3bc3727f-5260-4f4d-a829-963bcee6f22a · outbound
Safeguarding Large Language Models in Real-time with Tunable Safety-Performance Trade-offs LLM Censorship: A Machine Learning Challenge or a Computer Security Problem?
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a1b025e5-7b99-4884-840d-ea2f15c73eb4 · outbound
Safeguarding Large Language Models in Real-time with Tunable Safety-Performance Trade-offs Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ba8e6823-c1d7-4ac5-a38f-12ad31215c0d · outbound
Safeguarding Large Language Models in Real-time with Tunable Safety-Performance Trade-offs Malla: Demystifying Real-world Large Language Model Integrated Malicious Services
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4835ab57-1cb7-41c6-a964-c80b41c647e7 · outbound
Safeguarding Large Language Models in Real-time with Tunable Safety-Performance Trade-offs AutoDAN: Generating Stealthy Jailbreak Prompts on Aligned Large Language Models
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8fceeec4-0ecb-4777-8147-795effd35506 · outbound
Safeguarding Large Language Models in Real-time with Tunable Safety-Performance Trade-offs RoBERTa: A Robustly Optimized BERT Pretraining Approach
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 12a1b381-7c3f-483d-b292-82a9dfcb8018 · outbound
Safeguarding Large Language Models in Real-time with Tunable Safety-Performance Trade-offs arXiv preprint arXiv:2408.15625
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6119d039-e7b3-4081-8fb2-678139218fa9 · outbound
Safeguarding Large Language Models in Real-time with Tunable Safety-Performance Trade-offs Red Teaming Language Models with Language Models
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1b6ebc3f-f1d0-4aca-94ae-fcb85c3cac89 · outbound
Safeguarding Large Language Models in Real-time with Tunable Safety-Performance Trade-offs Bergeron: Combating Adversarial Attacks through a Conscience-Based Alignment Framework
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d1ccc78b-cced-467b-ae26-ea4e64fe8d4f · outbound
Safeguarding Large Language Models in Real-time with Tunable Safety-Performance Trade-offs Controllable Natural Language Generation with Contrastive Prefixes
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7a326ec1-0c5e-4501-a40c-99b3d892d5d8 · outbound
Safeguarding Large Language Models in Real-time with Tunable Safety-Performance Trade-offs In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 26473cd4-d6af-4c70-815c-bc37de93f660 · outbound
Safeguarding Large Language Models in Real-time with Tunable Safety-Performance Trade-offs Jailbreak Antidote: Runtime Safety-Utility Balance via Sparse Representation Adjustment in Large Language Models
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 37970b0b-1c64-4e4f-bf80-8bc2e3f569fa · outbound
Safeguarding Large Language Models in Real-time with Tunable Safety-Performance Trade-offs "Do Anything Now": Characterizing and Evaluating In-The-Wild Jailbreak Prompts on Large Language Models
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c2c08826-62ac-4dc9-a540-48f58aaabbc5 · outbound
Safeguarding Large Language Models in Real-time with Tunable Safety-Performance Trade-offs Self-Guard: Empower the LLM to Safeguard Itself
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d6ef1e05-0fd2-4d45-bd4d-a8d6458c5f61 · outbound
Safeguarding Large Language Models in Real-time with Tunable Safety-Performance Trade-offs Jailbreak and Guard Aligned Language Models with Only Few In-Context Demonstrations
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5ef4fb56-bdc6-405e-9fff-bfa24ccbd7bb · outbound
Safeguarding Large Language Models in Real-time with Tunable Safety-Performance Trade-offs In Findings of the Association for Computational Lin- guistics ACL 2024, pages 7432–7449
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation a76bef2b-afc8-45d7-b375-199041cb38a0 · outbound
Safeguarding Large Language Models in Real-time with Tunable Safety-Performance Trade-offs Low-Resource Languages Jailbreak GPT-4
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 85e48ec5-0dfa-49f5-84f6-0a6fea4bd398 · outbound
Safeguarding Large Language Models in Real-time with Tunable Safety-Performance Trade-offs Instruction-Following Evaluation for Large Language Models
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b8c638d9-7e46-47fb-a89c-5652165497e9 · outbound
Safeguarding Large Language Models in Real-time with Tunable Safety-Performance Trade-offs Universal and Transferable Adversarial Attacks on Aligned Language Models
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ff7fe4c1-f2c9-40e9-a15f-d1894817d866 · outbound
Safeguarding Large Language Models in Real-time with Tunable Safety-Performance Trade-offs A Framework for Real-time Safeguarding the Text Generation of Large Language Model
Reference 1996
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2b35992b-33b9-42ff-928b-382d8ce35a2e · outbound
Safeguarding Large Language Models in Real-time with Tunable Safety-Performance Trade-offs Machine Teaching: A New Paradigm for Building Machine Learning Systems
Reference 2017
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5b45da61-d38f-4f7a-8825-1bdee77aea47 · outbound
Safeguarding Large Language Models in Real-time with Tunable Safety-Performance Trade-offs Situating Sentence Embedders with Nearest Neighbor Overlap
Reference 2019
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 25bf91ad-1078-4808-8f63-e48ad7594a15 · outbound
Safeguarding Large Language Models in Real-time with Tunable Safety-Performance Trade-offs GeDi: Generative Discriminator Guided Sequence Generation
Reference 2020
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 94a17b93-d682-49d3-a001-16291b145f60 · outbound
Safeguarding Large Language Models in Real-time with Tunable Safety-Performance Trade-offs FUDGE: Controlled Text Generation With Future Discriminators
Reference 2021
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d264333e-f9dc-43d0-8b24-3d181803a16e · outbound
Safeguarding Large Language Models in Real-time with Tunable Safety-Performance Trade-offs Critic-Guided Decoding for Controlled Text Generation
Reference 2022
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 44ce380b-daee-45f6-aa11-b23be1c1c8f6 · outbound
Safeguarding Large Language Models in Real-time with Tunable Safety-Performance Trade-offs Detecting Language Model Attacks with Perplexity
Reference 2023
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bf947123-d390-486c-a39a-a8ba06f22841 · outbound
Safeguarding Large Language Models in Real-time with Tunable Safety-Performance Trade-offs In ICLR 2024 Workshop on Secure and Trustworthy Large Language Models
Reference 2024
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 88536e69-d799-494d-9d17-6eca02ab0ea3 · inbound
On Almost Surely Safe Alignment of Large Language Models at Inference-Time Safeguarding Large Language Models in Real-time with Tunable Safety-Performance Trade-offs
Reference 117
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b9e927c0-0ce1-4119-870b-44d6bb06c864 · inbound
The Future is Agentic: Definitions, Perspectives, and Open Challenges of Multi-Agent Recommender Systems Safeguarding Large Language Models in Real-time with Tunable Safety-Performance Trade-offs
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 105f2060-0c23-46f7-adee-753beda65c02 · inbound
CARE: Decoding Time Safety Alignment via Rollback and Introspection Intervention Safeguarding Large Language Models in Real-time with Tunable Safety-Performance Trade-offs
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.