Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-15T16:25:35.794271Z
Paper Citation Record · LEDGER
As of 15 August 2026, this Paper Citation Record lists 91 of 91 outbound references and 1 inbound Pith citation observation for arXiv:2509.08000.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-15T16:25:35.794271Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-06-27T03:24:24.714121Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z
91 of 91 outbound references displayed
External citation measurements
0
arxiv_reference, observed 2026-08-05T02:28:24.338817Z
Observation ad25e298-6573-45c4-85c6-f3496b163bae · outbound
AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMs , " * write output.state after.block = add.period write newline
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation aa8a2ef4-2be3-40f2-a0b0-5e1d45d5e73c · outbound
AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMs write newline
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5cb1a105-d857-45f5-bd98-f79924b9a913 · outbound
AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMs Yi: Open Foundation Models by 01.AI
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 63e14614-66d4-4149-b35d-a0c099c42125 · outbound
AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMs The Falcon Series of Open Language Models
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 00948fef-d163-4498-bb1e-14fb8d997059 · outbound
AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMs E.; Hubinger, E.; Bai, Y.; Bricken, T.; Maxwell, T.; Schiefer, N.; Sully, J.; Tamkin, A.; Lanham, T.; Nguyen, K.; Korbak, T.; Kaplan, J.; Ganguli, D.; Bowman, S
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8f4c4d3c-4ac2-414d-a932-56dd9d53052f · outbound
AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMs Refusal in Language Models Is Mediated by a Single Direction
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation af39a24f-0b0c-427b-9445-15d5df9c7f85 · outbound
AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMs Do Anything Now
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a601b59f-ea31-4e22-a836-625c583ea314 · outbound
AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMs Unresolved cited work
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation da90c372-54be-4d7a-a2d3-25ba07ff3003 · outbound
AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMs A.; and Hill, E
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 19928a12-01f7-45d0-af60-33a7fee39670 · outbound
AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMs Unresolved cited work
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9179db64-dab0-43ef-b687-c3e2d825e606 · outbound
AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMs Unresolved cited work
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 37912514-015c-4863-89a5-7890860e2bbb · outbound
AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMs J.; and Wong, E
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d4017d78-8c14-4e7f-a732-46170f55fb0a · outbound
AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMs Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4859decf-14a6-450c-af55-fa9e40cf246a · outbound
AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMs Unresolved cited work
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3d6f2e35-02f0-4520-9bab-70d7fe82028a · outbound
AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMs Training Verifiers to Solve Math Word Problems
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c9f372f3-b706-4e65-b29e-3c151c9d96a2 · outbound
AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMs Multi-Head Attention: Collaborate Instead of Concatenate
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 029c671a-408e-49ce-810b-7cc95ca2f8b1 · outbound
AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMs DeepSeek-V3 Technical Report
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e86dd41e-677f-408d-91a9-fb22a54ff9e6 · outbound
AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMs Unresolved cited work
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3647b421-971a-4cfb-9afb-4c2ce60a1fb3 · outbound
AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMs C.; Allen, E
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 40ca8320-99af-4c93-ae36-9989f479325e · outbound
AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMs Unresolved cited work
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c35178b1-326b-458b-af3d-87ab63237d04 · outbound
AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMs Unresolved cited work
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 45c5d733-ba60-419d-916d-8e2335e4d032 · outbound
AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMs Unresolved cited work
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 4c40d6e0-a570-4127-b746-4696d7e71c9f · outbound
AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMs FigStep: Jailbreaking Large Vision-Language Models via Typographic Visual Prompts
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9009f364-c983-445c-afc9-abd99ed4df17 · outbound
AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMs The Llama 3 Herd of Models
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 31fbdf9e-ce32-419c-bb86-c8bc84d5287a · outbound
AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMs Unresolved cited work
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 4114cf85-cd5d-4040-aece-c1df0a48ff81 · outbound
AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMs Aloe: A Family of Fine-tuned Open Healthcare LLMs
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d062791c-dbb2-4b6b-b6a7-21ede5e85ef2 · outbound
AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMs Measuring Massive Multitask Language Understanding
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 48e79834-71bf-48de-b8f7-b22018b76567 · outbound
AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMs Measuring Mathematical Problem Solving With the MATH Dataset
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9980664b-b9ad-4173-a854-2838d6d1b1b1 · outbound
AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMs Unnatural Instructions: Tuning Language Models with (Almost) No Human Labor
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bc112cc0-46f6-4001-9b2d-bd1e020d062c · outbound
AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMs LoRA: Low-Rank Adaptation of Large Language Models
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 468945ff-c394-4a8b-aae3-214d259cdeb7 · outbound
AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMs Booster: Tackling Harmful Fine-tuning for Large Language Models via Attenuating Harmful Perturbation
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f5f1cb5b-891b-44da-950b-d6cac8a6d69c · outbound
AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMs Virus: Harmful Fine-tuning Attack for Large Language Models Bypassing Guardrail Moderation
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 18be176f-584e-40d0-86c2-d944de26f4f9 · outbound
AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMs Vaccine: Perturbation-aware Alignment for Large Language Models against Harmful Fine-tuning Attack
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d04578d6-677e-4791-b157-4fac8425e1a4 · outbound
AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMs Unresolved cited work
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 8dbbecfb-e8b7-4477-a8ea-9051a1b85c0a · outbound
AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMs Large Language Model (LLM) for Software Security: Code Analysis, Malware Analysis, Reverse Engineering
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 10c4157a-e12b-43ce-80dc-ec8054f35342 · outbound
AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMs Unresolved cited work
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 96d1e257-d944-45cf-9a65-8e2a6c5847ff · outbound
AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMs Mistral 7B
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 70274a89-1af6-4411-8775-a947129265ee · outbound
AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMs Unresolved cited work
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation e6ae978d-4772-4f3f-a900-1d9ec49f6752 · outbound
AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMs LLM Embeddings for Deep Learning on Tabular Data
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3dde08bb-842c-49fb-869d-762eabf1e434 · outbound
AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMs Robust Distortion-free Watermarks for Language Models
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 46e0b858-00b3-4787-9db8-d7fe7ad30e9e · outbound
AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMs Unresolved cited work
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 926a9815-7dac-4d80-a119-e24de3292678 · outbound
AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMs Multi-step Jailbreaking Privacy Attacks on ChatGPT
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c75d8f18-27c6-478c-aa7e-ace93bf76c7e · outbound
AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMs The WMDP Benchmark: Measuring and Reducing Malicious Use With Unlearning
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b91a5ea8-9ede-4d8a-87b6-4166cec6f27f · outbound
AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMs TeleLoRA: Teleporting Model-Specific Alignment Across LLMs
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 2d16c294-5216-4998-9779-f7dbae85971a · outbound
AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMs Unresolved cited work
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 0e368ff1-ba39-4568-88d8-30414caf00aa · outbound
AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMs DeepInception: Hypnotize Large Language Model to Be Jailbreaker
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c47c0b94-a64a-4edc-8d32-104d6ebbdb54 · outbound
AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMs SHIELD: Evaluation and Defense Strategies for Copyright Compliance in LLM Text Generation
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ca700848-885d-4c03-8a6f-32e5254d0726 · outbound
AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMs AutoDAN: Generating Stealthy Jailbreak Prompts on Aligned Large Language Models
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 98430e03-cf23-443b-9007-c53b22089ba1 · outbound
AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMs HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust Refusal
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ca845d39-ef48-4bbc-9588-f5a40d6a5c14 · outbound
AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMs R.; and Papernot, N
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 833a1260-eedb-4242-bd8a-bccd6a290aab · outbound
AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMs Topic-Based Watermarks for Large Language Models
Reference 57
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation eb29432f-9822-466a-b1fa-8fa53f9fb69f · outbound
AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMs J.; Hassani, H.; Robey, A.; and Wong, E
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 0edacc3f-f76f-4b44-b7a7-038ce3b123f5 · outbound
AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMs In-Context Unlearning: Language Models as Few Shot Unlearners
Reference 59
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c3ff32b1-5882-4710-a2ed-0495bdeabd70 · outbound
AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMs Unresolved cited work
Reference 60
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation ddceaa23-7b54-4660-a7df-0af32dfa13e8 · outbound
AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMs LLM Defenses Are Not Robust to Multi-Turn Human Jailbreaks Yet
Reference 61
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 53b212db-a8ec-46fa-84a7-f7fea271f2ea · outbound
AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMs Direct Preference Optimization: Your Language Model is Secretly a Reward Model
Reference 62
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3ce97506-2eff-4deb-b347-a0fcd6367b25 · outbound
AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMs Unresolved cited work
Reference 63
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation e76398f4-5be7-4001-bc36-72747e6bb4ad · outbound
AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMs XSTest: A Test Suite for Identifying Exaggerated Safety Behaviours in Large Language Models
Reference 64
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0cf543b4-961c-4339-a167-43334dd940df · outbound
AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMs Unresolved cited work
Reference 65
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 529d0413-b7ce-4b21-969d-7682e60c144b · outbound
AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMs Language Models Are Greedy Reasoners: A Systematic Formal Analysis of Chain-of-Thought
Reference 66
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0054103c-e053-4a9b-87c5-b95df03abe64 · outbound
AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMs Scalable and Transferable Black-Box Jailbreaks for Language Models via Persona Modulation
Reference 67
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 399fe1e5-aac0-4f27-a586-77d9448e3b04 · outbound
AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMs Unresolved cited work
Reference 68
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 36314a39-cde8-4e78-9130-60869cbd74c4 · outbound
AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMs Survey of Vulnerabilities in Large Language Models Revealed by Adversarial Attacks
Reference 69
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9ac3813f-9ba9-4c50-a3b5-403f03f45be1 · outbound
AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMs Do Anything Now
Reference 70
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation d1af0da0-2d5e-4a04-80e3-98009e17e73c · outbound
AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMs "Do Anything Now": Characterizing and Evaluating In-The-Wild Jailbreak Prompts on Large Language Models
Reference 71
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0df81bd0-f189-401c-b4da-e05e884a8bfa · outbound
AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMs Counterfactual Explainable Incremental Prompt Attack Analysis on Large Language Models
Reference 72
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 6d520293-7269-4cab-88c1-5d92ee86b72c · outbound
AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMs A StrongREJECT for Empty Jailbreaks
Reference 73
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b5dffa9a-35b4-4ccc-af4a-d3635da1ee57 · outbound
AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMs Tamper-Resistant Safeguards for Open-Weight LLMs
Reference 74
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8e6337e5-00e5-4a69-9bae-f25b6ef9dfa9 · outbound
AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMs Gemma 3 Technical Report
Reference 75
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9af2336c-692a-4b0d-a3fa-f783b16507e6 · outbound
AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMs Unresolved cited work
Reference 76
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5655b3f3-d956-4b39-99a4-a1fa20e609ec · outbound
AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMs Prompt, Divide, and Conquer: Bypassing Large Language Model Safety Filters via Segmented and Distributed Prompt Processing
Reference 77
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 874015b7-cb02-40d5-924a-cdcdf96b9abf · outbound
AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMs Do-Not-Answer: A Dataset for Evaluating Safeguards in LLMs
Reference 78
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 51df8078-f20f-4b92-8d42-72156df20f57 · outbound
AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMs Jailbroken: How Does LLM Safety Training Fail?
Reference 80
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 54f0fe5f-a4cd-4e81-974e-d7c8efcbbd7b · outbound
AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMs SMILES-Prompting: A Novel Approach to LLM Jailbreak Attacks in Chemical Synthesis
Reference 81
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e8441d64-02c5-48ec-880f-9169d9c2e1e5 · outbound
AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMs Unresolved cited work
Reference 82
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation b33b4b48-4e00-4557-b1d1-80b2dd061767 · outbound
AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMs Unresolved cited work
Reference 83
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 7744cce7-13fd-4e4c-a439-05382bd6f77b · outbound
AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMs RedAgent: Red Teaming Large Language Models with Context-aware Autonomous Language Agent
Reference 84
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6600dd76-3429-4677-9637-7463b1b07cf3 · outbound
AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMs Unresolved cited work
Reference 85
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 602c5ab8-683e-42e2-b488-b0ba8192c37d · outbound
AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMs Qwen3 Technical Report
Reference 86
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6737df63-0dd1-4039-b648-154aac2f6535 · outbound
AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMs Jigsaw Puzzles: Splitting Harmful Questions to Jailbreak Large Language Models
Reference 87
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7bd9f140-c927-4bba-9625-210afbdebf6d · outbound
AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMs The Dark Side of Trust: Authority Citation-Driven Jailbreak Attacks on Large Language Models
Reference 88
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a969486c-572c-44f7-aaab-60fbad43883b · outbound
AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMs Low-Resource Languages Jailbreak GPT-4
Reference 89
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6b76a273-6f9b-4777-8418-4253fb003343 · outbound
AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMs Position: Editing Large Language Models Poses Serious Safety Risks
Reference 90
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c7c025be-ad8a-434b-ab31-d3acc0e091cb · outbound
AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMs Don't Listen To Me: Understanding and Exploring Jailbreak Prompts of Large Language Models
Reference 92
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2d1cdad1-23db-49bf-973b-2304ef7d8878 · outbound
AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMs HellaSwag: Can a Machine Really Finish Your Sentence?
Reference 93
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f7125f4b-9524-4f2f-bb51-c7472aa3e70c · outbound
AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMs How Johnny Can Persuade LLMs to Jailbreak Them: Rethinking Persuasion to Challenge AI Safety by Humanizing LLMs
Reference 95
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 51ede104-63ed-4a65-95df-a7068fb2edba · outbound
AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMs LLM Embeddings Improve Test-time Adaptation to Tabular $Y|X$-Shifts
Reference 96
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 7e84904e-ce8f-4cf2-b123-5d300a9c3e9d · outbound
AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMs Unresolved cited work
Reference 97
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 7d03a866-8998-4115-942f-ca7cbea3c008 · outbound
AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMs Making Them Ask and Answer: Jailbreaking Large Language Models in Few Queries via Disguise and Reconstruction
Reference 98
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2511efdb-5ff7-41fc-acc9-8bd3f5727caf · outbound
AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMs LIMA: Less Is More for Alignment
Reference 99
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 82f303b1-65e0-4869-8233-fc05e3500b2e · outbound
AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMs Aya Model: An Instruction Finetuned Open-Access Multilingual Language Model
Reference 100
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9d9a0935-e55e-49d7-b954-564f6fade08c · inbound
Do Activation Monitors Survive Model Updates? Benchmarking, Predicting, and Repairing Activation-Monitor Staleness AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMs
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.