Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T12:05:44.625951Z
Paper Citation Record · LEDGER
As of 9 August 2026, this Paper Citation Record lists 33 of 33 outbound references and 1 inbound Pith citation observation for arXiv:2507.22160.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T12:05:44.625951Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-03T14:24:24.267269Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-08-05T02:28:24.338817Z
33 of 33 outbound references displayed
External citation measurements
0
pith, observed 2026-08-05T02:28:24.338817Z
Observation 21742424-0228-49b6-b944-22f39c585070 · outbound
Strategic Deflection: Defending LLMs from Logit Manipulation Training language models to follow instructions with human feedback
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 16a914b2-ff48-422c-91e7-ace163ab7bcf · outbound
Strategic Deflection: Defending LLMs from Logit Manipulation Deep reinforcement learning from human preferences
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation dfa9d83a-feca-46da-a357-460ec701c679 · outbound
Strategic Deflection: Defending LLMs from Logit Manipulation Mistral 7B
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f09ff4ca-dc03-46de-b451-fd2cba7fc169 · outbound
Strategic Deflection: Defending LLMs from Logit Manipulation The llama 3 herd of models
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8b56acc9-7aa6-4136-ba2a-9940a912e427 · outbound
Strategic Deflection: Defending LLMs from Logit Manipulation On large language models’ resilience to coercive interrogation
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 6ab8a4e0-999b-4b93-a1ca-820f02528e21 · outbound
Strategic Deflection: Defending LLMs from Logit Manipulation Universal and Transferable Adversarial Attacks on Aligned Language Models
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5d7d4768-5c38-442e-900d-7453e798e9ea · outbound
Strategic Deflection: Defending LLMs from Logit Manipulation Jailbreak open-sourced large language models via enforced decoding
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 99440afc-c6b0-4c6b-9eb2-c10f3cdb3c0e · outbound
Strategic Deflection: Defending LLMs from Logit Manipulation Safety alignment should be made more than just a few tokens deep
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 180b2f18-aaf7-4ad0-86ef-3d9682ca6f8c · outbound
Strategic Deflection: Defending LLMs from Logit Manipulation Jailbroken: How does llm safety training fail? In Advances in Neural Information Processing Systems , volume 36, pages 80079–80110, 2023
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation e838db4e-e370-4e8c-b125-56c08475e087 · outbound
Strategic Deflection: Defending LLMs from Logit Manipulation do anything now
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 4fd0b2e3-1ba8-43d1-9d0e-8a75fcc6e530 · outbound
Strategic Deflection: Defending LLMs from Logit Manipulation A Cross-Language Investigation into Jailbreak Attacks in Large Language Models
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6b12d06e-437a-4d7b-9924-6203a60c4777 · outbound
Strategic Deflection: Defending LLMs from Logit Manipulation Protecting your llms with information bottleneck
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 959e6f3d-fa51-4b4f-ab98-39129bcc81b5 · outbound
Strategic Deflection: Defending LLMs from Logit Manipulation Llama 2: Open Foundation and Fine-Tuned Chat Models
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7f96471a-2d22-4819-acaf-8b43d1aebc20 · outbound
Strategic Deflection: Defending LLMs from Logit Manipulation Jailbreaking black box large language models in twenty queries
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation ec7747ca-aa95-40b9-829b-cb04fb8f2627 · outbound
Strategic Deflection: Defending LLMs from Logit Manipulation Fuzzllm: A novel and universal fuzzing framework for proactively discovering jailbreak vulnerabilities in large language models
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation df81673c-0842-46ba-859d-37721d4b7f7d · outbound
Strategic Deflection: Defending LLMs from Logit Manipulation Catastrophic jailbreak of open-source LLMs via exploiting generation
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 3c873539-fcdf-4f0e-ac3e-142987dd9777 · outbound
Strategic Deflection: Defending LLMs from Logit Manipulation Baseline Defenses for Adversarial Attacks Against Aligned Language Models
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 15f793b9-b949-4dce-9e30-e3e82aa31e24 · outbound
Strategic Deflection: Defending LLMs from Logit Manipulation Certifying LLM Safety against Adversarial Prompting
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fb9a6b82-8580-41da-95fa-7ed979b2a608 · outbound
Strategic Deflection: Defending LLMs from Logit Manipulation Detecting Language Model Attacks with Perplexity
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7f9a2905-5fa5-454b-8b4f-e8392af0d7b5 · outbound
Strategic Deflection: Defending LLMs from Logit Manipulation Safety at Scale: A Comprehensive Survey of Large Model and Agent Safety
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5f17508b-2095-457f-8ce1-88baba66d487 · outbound
Strategic Deflection: Defending LLMs from Logit Manipulation Robust safety classifier against jailbreaking attacks: Adversarial prompt shield
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation cf1f2e26-6ed2-4ed9-ae91-db14b03b78ea · outbound
Strategic Deflection: Defending LLMs from Logit Manipulation Jailbreaking leading safety-aligned LLMs with simple adaptive attacks
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 96d42cbc-56fc-4e49-9716-d763f85de590 · outbound
Strategic Deflection: Defending LLMs from Logit Manipulation Contrastive preference optimization: pushing the boundaries of llm performance in machine translation
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 5a10df76-0587-463d-ae29-e55144495ca9 · outbound
Strategic Deflection: Defending LLMs from Logit Manipulation Direct preference optimization: Your language model is secretly a reward model
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 30688163-6060-4951-9142-cd1519a81b63 · outbound
Strategic Deflection: Defending LLMs from Logit Manipulation GPT-4o System Card
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0ccbce3f-2f9f-4247-abac-6a0be005e2f2 · outbound
Strategic Deflection: Defending LLMs from Logit Manipulation LoRA: Low-rank adaptation of large language models
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation bc351a75-c6fb-427e-acbc-49c9ac3f7208 · outbound
Strategic Deflection: Defending LLMs from Logit Manipulation Trl: Transformer reinforcement learning
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 873c020d-26d4-407f-8a25-49f083c58d78 · outbound
Strategic Deflection: Defending LLMs from Logit Manipulation Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0ae32380-0592-472e-9f90-fdee11c65965 · outbound
Strategic Deflection: Defending LLMs from Logit Manipulation tinybench- marks: evaluating llms with fewer examples
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 7681c4da-35cc-4c17-88e6-1bca9d12be53 · outbound
Strategic Deflection: Defending LLMs from Logit Manipulation Measuring massive multitask language understanding
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation f5ba09d9-18eb-47c9-8748-12d938f62721 · outbound
Strategic Deflection: Defending LLMs from Logit Manipulation Hellaswag: Can a machine really finish your sentence? In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics , pages 4791–4800, 2019
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation e1449f35-7554-4725-8fee-7f70f47afc52 · outbound
Strategic Deflection: Defending LLMs from Logit Manipulation Truthfulqa: Measuring how models mimic human falsehoods
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation a8c10808-bfa8-4ae3-b582-4d13098932b0 · outbound
Strategic Deflection: Defending LLMs from Logit Manipulation Training Verifiers to Solve Math Word Problems
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6555fb52-5a81-4d7c-a31d-47c9eba147db · inbound
Safety Alignment of LMs via Non-cooperative Games Strategic Deflection: Defending LLMs from Logit Manipulation
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.