Pith. sign in

Paper Citation Record · LEDGER

AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMs

As of 15 August 2026, this Paper Citation Record lists 91 of 91 outbound references and 1 inbound Pith citation observation for arXiv:2509.08000.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2509.08000 v1

Coverage vector

measured 91 of 91 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T16:25:35.794271Z

measured 92 of 92 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-06-27T03:24:24.714121Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

91 of 91 outbound references displayed

  • verified exact4
  • verified fuzzy2
  • unresolved85
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

0
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation ad25e298-6573-45c4-85c6-f3496b163bae · outbound

This paper cites , " * write output.state after.block = add.period write newline.

AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMs , " * write output.state after.block = add.period write newline

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-15T16:25:35.382342Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T16:25:35.382342Z digest=sha256:0e8ae91da622845e13076ea984f9267ea7005e3ff0e048d92211939532d279f3

Observation aa8a2ef4-2be3-40f2-a0b0-5e1d45d5e73c · outbound

This paper cites write newline.

AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMs write newline

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-15T16:25:35.387959Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T16:25:35.387959Z digest=sha256:e6733f872c7fddb5fde1e4aa03db63b8aa93228dc90527f288413c8e0cdf03c6

Observation 5cb1a105-d857-45f5-bd98-f79924b9a913 · outbound

This paper cites Yi: Open Foundation Models by 01.AI.

AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMs Yi: Open Foundation Models by 01.AI

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-15T16:25:35.392417Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T16:25:35.392417Z digest=sha256:d4c44749cd8aad3ecba3100eb065127eda31bdcc22efb2348cf19ab2f24e1c3f

Observation 63e14614-66d4-4149-b35d-a0c099c42125 · outbound

This paper cites The Falcon Series of Open Language Models.

AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMs The Falcon Series of Open Language Models

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-15T16:25:35.397446Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T16:25:35.397446Z digest=sha256:59e20888aea4c5d54d99c910b3beeddf941f1f3d7e60de6544b15a063a68bfe0

Observation 00948fef-d163-4498-bb1e-14fb8d997059 · outbound

This paper cites E.; Hubinger, E.; Bai, Y.; Bricken, T.; Maxwell, T.; Schiefer, N.; Sully, J.; Tamkin, A.; Lanham, T.; Nguyen, K.; Korbak, T.; Kaplan, J.; Ganguli, D.; Bowman, S.

AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMs E.; Hubinger, E.; Bai, Y.; Bricken, T.; Maxwell, T.; Schiefer, N.; Sully, J.; Tamkin, A.; Lanham, T.; Nguyen, K.; Korbak, T.; Kaplan, J.; Ganguli, D.; Bowman, S

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-15T16:25:35.401443Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T16:25:35.401443Z digest=sha256:5654b1569d45a44388af45ce38f08d06808597636a8790c63ade19be890dc5f2

Observation 8f4c4d3c-4ac2-414d-a932-56dd9d53052f · outbound

This paper cites Refusal in Language Models Is Mediated by a Single Direction.

AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMs Refusal in Language Models Is Mediated by a Single Direction

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-15T16:25:35.405667Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T16:25:35.405667Z digest=sha256:13089604b138dd19a1b0cae555d9f978990ba104fb0eb5d7fdb78b25073258f4

Observation af39a24f-0b0c-427b-9445-15d5df9c7f85 · outbound

This paper cites Do Anything Now.

AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMs Do Anything Now

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-15T16:25:35.409950Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T16:25:35.409950Z digest=sha256:9abbf4729a1e6ed288f38152be1bda7d2c4d10a9d65d8830cbc7de6199366ced

Observation a601b59f-ea31-4e22-a836-625c583ea314 · outbound

This paper cites an unresolved cited work.

AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMs Unresolved cited work

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-15T16:25:35.413726Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T16:25:35.413726Z digest=sha256:4efe7101d4cd7537594282cfa2605877dc208ab82e35bef757fbd16370bd27f1

Observation da90c372-54be-4d7a-a2d3-25ba07ff3003 · outbound

This paper cites A.; and Hill, E.

AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMs A.; and Hill, E

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-15T16:25:35.418714Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T16:25:35.418714Z digest=sha256:bf046f3ccb05de8c8d54ecf30a2f97f381e5a0bc2dcabeb4f7f55e1c63bd6476

Observation 19928a12-01f7-45d0-af60-33a7fee39670 · outbound

This paper cites an unresolved cited work.

AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMs Unresolved cited work

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-15T16:25:35.423138Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T16:25:35.423138Z digest=sha256:f0cfe98aa86fa97dd5480315bb33a1beecdf54780fc4b2c3e51adf3402d3a3fb

Observation 9179db64-dab0-43ef-b687-c3e2d825e606 · outbound

This paper cites an unresolved cited work.

AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMs Unresolved cited work

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-15T16:25:35.429822Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T16:25:35.429822Z digest=sha256:a6e120f8c86f8497591f21f007dd926e5e7741cc01e2d9ca118003aee5bdee33

Observation 37912514-015c-4863-89a5-7890860e2bbb · outbound

This paper cites J.; and Wong, E.

AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMs J.; and Wong, E

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-15T16:25:35.438338Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T16:25:35.438338Z digest=sha256:04ba440012c4677449a32909910c56d0bdda086f3d417a137f9815b91f9ded2e

Observation d4017d78-8c14-4e7f-a732-46170f55fb0a · outbound

This paper cites Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities.

AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMs Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-15T16:25:35.442838Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T16:25:35.442838Z digest=sha256:56ed48112be0b47f99716e9e3c1a6c03f83bbe8b39d257268e0dd6a3d3df6421

Observation 4859decf-14a6-450c-af55-fa9e40cf246a · outbound

This paper cites an unresolved cited work.

AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMs Unresolved cited work

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-15T16:25:35.447097Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T16:25:35.447097Z digest=sha256:1a161e9f913cef9537d843aa4337db6f540f99fe5d1968152bce651ad36584b4

Observation 3d6f2e35-02f0-4520-9bab-70d7fe82028a · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMs Training Verifiers to Solve Math Word Problems

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-15T16:25:35.450756Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T16:25:35.450756Z digest=sha256:27a11730e98143c48fdd30618642bcfe8cfde4328d2afa826ded1926b58df5f0

Observation c9f372f3-b706-4e65-b29e-3c151c9d96a2 · outbound

This paper cites Multi-Head Attention: Collaborate Instead of Concatenate.

AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMs Multi-Head Attention: Collaborate Instead of Concatenate

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-15T16:25:35.454813Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T16:25:35.454813Z digest=sha256:5df597d5e7a3d1e961d61852d32898996ac731a2f065dffe6e696b495f671208

Observation 029c671a-408e-49ce-810b-7cc95ca2f8b1 · outbound

This paper cites DeepSeek-V3 Technical Report.

AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMs DeepSeek-V3 Technical Report

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-15T16:25:35.459056Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T16:25:35.459056Z digest=sha256:94b58731cb5a0a80879d5f74aea06e619447384932d43ad6658f2d6d67828434

Observation e86dd41e-677f-408d-91a9-fb22a54ff9e6 · outbound

This paper cites an unresolved cited work.

AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMs Unresolved cited work

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-15T16:25:35.463314Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T16:25:35.463314Z digest=sha256:bfd18acd43422378c008211e08ea39bde7f72d35dde6dee5965f347a8e7b0391

Observation 3647b421-971a-4cfb-9afb-4c2ce60a1fb3 · outbound

This paper cites C.; Allen, E.

AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMs C.; Allen, E

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-15T16:25:35.466791Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T16:25:35.466791Z digest=sha256:f905d1d381bc8de48b6c310694ba95277fee91adeff083e810d888ee704ddc8f

Observation 40ca8320-99af-4c93-ae36-9989f479325e · outbound

This paper cites an unresolved cited work.

AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMs Unresolved cited work

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-15T16:25:35.470899Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T16:25:35.470899Z digest=sha256:abb57d19507da07fc27692d62c60ede52175cd018560e4e26bf5b3900108c041

Observation c35178b1-326b-458b-af3d-87ab63237d04 · outbound

This paper cites an unresolved cited work.

AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMs Unresolved cited work

Reference 21

Resolution
unresolved
raw_fallback, observed 2026-08-15T16:25:37.246678Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-15T16:25:35.475467Z digest=sha256:dd8959a991bf3ff4ec42185fc214d35787ee3f18de5b2977f2b7710c39380210

Observation 45c5d733-ba60-419d-916d-8e2335e4d032 · outbound

This paper cites an unresolved cited work.

AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMs Unresolved cited work

Reference 22

Resolution
unresolved
raw_fallback, observed 2026-08-15T16:25:37.231653Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-15T16:25:35.479137Z digest=sha256:7daefd2f7aa9bade33c1f3aafca0d96821b685fe8dc6b86d61a009d1100f6525

Observation 4c40d6e0-a570-4127-b746-4696d7e71c9f · outbound

This paper cites FigStep: Jailbreaking Large Vision-Language Models via Typographic Visual Prompts.

AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMs FigStep: Jailbreaking Large Vision-Language Models via Typographic Visual Prompts

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-15T16:25:35.482899Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T16:25:35.482899Z digest=sha256:f0375fad352b95a9f4e8ee695916a96dec267bdd108ed2b7106931bae1ca483f

Observation 9009f364-c983-445c-afc9-abd99ed4df17 · outbound

This paper cites The Llama 3 Herd of Models.

AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMs The Llama 3 Herd of Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-15T16:25:35.486772Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T16:25:35.486772Z digest=sha256:c4964d766a4106d6e682035c55318d49befbd6272e6deb97527f1a204256ff8f

Observation 31fbdf9e-ce32-419c-bb86-c8bc84d5287a · outbound

This paper cites an unresolved cited work.

AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMs Unresolved cited work

Reference 25

Resolution
unresolved
raw_fallback, observed 2026-08-15T16:25:37.218168Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-15T16:25:35.491734Z digest=sha256:788f1a4550de23798c2c127ac99135b81316f55499e4202c37f88cab6ea7614b

Observation 4114cf85-cd5d-4040-aece-c1df0a48ff81 · outbound

This paper cites Aloe: A Family of Fine-tuned Open Healthcare LLMs.

AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMs Aloe: A Family of Fine-tuned Open Healthcare LLMs

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-15T16:25:35.495639Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T16:25:35.495639Z digest=sha256:241330ee567a115330495d4b8747b59670e90ba9b4f964a6da504102613c24c4

Observation d062791c-dbb2-4b6b-b6a7-21ede5e85ef2 · outbound

This paper cites Measuring Massive Multitask Language Understanding.

AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMs Measuring Massive Multitask Language Understanding

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-15T16:25:35.500101Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T16:25:35.500101Z digest=sha256:1160906a27352b0f2812446a400ae4947375be5adab3a69639d9f5316eae2e2c

Observation 48e79834-71bf-48de-b8f7-b22018b76567 · outbound

This paper cites Measuring Mathematical Problem Solving With the MATH Dataset.

AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMs Measuring Mathematical Problem Solving With the MATH Dataset

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-15T16:25:35.504361Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T16:25:35.504361Z digest=sha256:e26202864dbf2c91b456abb355dc0bf98d1fafaeecca0892b7ee73d121b3eab2

Observation 9980664b-b9ad-4173-a854-2838d6d1b1b1 · outbound

This paper cites Unnatural Instructions: Tuning Language Models with (Almost) No Human Labor.

AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMs Unnatural Instructions: Tuning Language Models with (Almost) No Human Labor

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-15T16:25:35.508322Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T16:25:35.508322Z digest=sha256:9d2fb371f9a4dec66b58d63a0f80e526ee7a4fa79b73aceb154b958a46e097a8

Observation bc112cc0-46f6-4001-9b2d-bd1e020d062c · outbound

This paper cites LoRA: Low-Rank Adaptation of Large Language Models.

AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMs LoRA: Low-Rank Adaptation of Large Language Models

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-15T16:25:35.512022Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T16:25:35.512022Z digest=sha256:fabf4cade8d70e696b68c9693d8145e2c285bfe08c06c18f8c31cd2dd1b43766

Observation 468945ff-c394-4a8b-aae3-214d259cdeb7 · outbound

This paper cites Booster: Tackling Harmful Fine-tuning for Large Language Models via Attenuating Harmful Perturbation.

AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMs Booster: Tackling Harmful Fine-tuning for Large Language Models via Attenuating Harmful Perturbation

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-15T16:25:35.520496Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T16:25:35.520496Z digest=sha256:5c254fce7b091682e1b8487aace7f412b2df2193e514bfc27b162d2d53e1c052

Observation f5f1cb5b-891b-44da-950b-d6cac8a6d69c · outbound

This paper cites Virus: Harmful Fine-tuning Attack for Large Language Models Bypassing Guardrail Moderation.

AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMs Virus: Harmful Fine-tuning Attack for Large Language Models Bypassing Guardrail Moderation

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-15T16:25:35.524521Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T16:25:35.524521Z digest=sha256:6386b4f5d75614179823330105c13b289019a3eae75ee0f609a03a60831cd78e

Observation 18be176f-584e-40d0-86c2-d944de26f4f9 · outbound

This paper cites Vaccine: Perturbation-aware Alignment for Large Language Models against Harmful Fine-tuning Attack.

AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMs Vaccine: Perturbation-aware Alignment for Large Language Models against Harmful Fine-tuning Attack

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-15T16:25:35.533528Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T16:25:35.533528Z digest=sha256:5c3db7f2125f14f0aed3a42e41ea40dcbefa356eba8d108f8ed127dfb8ab5e4b

Observation d04578d6-677e-4791-b157-4fac8425e1a4 · outbound

This paper cites an unresolved cited work.

AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMs Unresolved cited work

Reference 36

Resolution
unresolved
raw_fallback, observed 2026-08-15T16:25:37.205244Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-15T16:25:35.537145Z digest=sha256:489729dc56a9b7f9ddcc01e61f6adca099f4be38b81c28e78bb579bbc59c5d26

Observation 8dbbecfb-e8b7-4477-a8ea-9051a1b85c0a · outbound

This paper cites Large Language Model (LLM) for Software Security: Code Analysis, Malware Analysis, Reverse Engineering.

AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMs Large Language Model (LLM) for Software Security: Code Analysis, Malware Analysis, Reverse Engineering

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-15T16:25:35.540991Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T16:25:35.540991Z digest=sha256:c4318eb98d9e9c209d00c86bc5d5e4d1e999a812f8c1910f7ed107bfa19e7641

Observation 10c4157a-e12b-43ce-80dc-ec8054f35342 · outbound

This paper cites an unresolved cited work.

AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMs Unresolved cited work

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-15T16:25:35.545273Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T16:25:35.545273Z digest=sha256:69ae59bd60b0cb3a5885ce8654897860f2706e4a65eb7c5a0fcd5f5013a0c008

Observation 96d1e257-d944-45cf-9a65-8e2a6c5847ff · outbound

This paper cites Mistral 7B.

AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMs Mistral 7B

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-15T16:25:35.548883Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T16:25:35.548883Z digest=sha256:09567289fd4826fe16a2f536ba985282a3540f50bc550d998de3ba6d6ba96cf2

Observation 70274a89-1af6-4411-8775-a947129265ee · outbound

This paper cites an unresolved cited work.

AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMs Unresolved cited work

Reference 40

Resolution
unresolved
raw_fallback, observed 2026-08-15T16:25:37.192162Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-15T16:25:35.553625Z digest=sha256:a1964c44418ac79dc40120ab6ba2ff2b470749d03d7176b2279acd6f21c7b2ca

Observation e6ae978d-4772-4f3f-a900-1d9ec49f6752 · outbound

This paper cites LLM Embeddings for Deep Learning on Tabular Data.

AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMs LLM Embeddings for Deep Learning on Tabular Data

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-15T16:25:35.558732Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T16:25:35.558732Z digest=sha256:e3f9888f9bce84497d2172e693a456584c0c4d96260bade52430cf6625f23914

Observation 3dde08bb-842c-49fb-869d-762eabf1e434 · outbound

This paper cites Robust Distortion-free Watermarks for Language Models.

AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMs Robust Distortion-free Watermarks for Language Models

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-15T16:25:35.562902Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T16:25:35.562902Z digest=sha256:65363de632f2c6737df1a0159eb51f0af0c7f876385eb2687d6555a157e71019

Observation 46e0b858-00b3-4787-9db8-d7fe7ad30e9e · outbound

This paper cites an unresolved cited work.

AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMs Unresolved cited work

Reference 43

Resolution
unresolved
raw_fallback, observed 2026-08-15T16:25:37.178690Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-15T16:25:35.566932Z digest=sha256:d33b74199e04690ea095c3899926e124b44caa8ec2863b5b3d05f633f5de2784

Observation 926a9815-7dac-4d80-a119-e24de3292678 · outbound

This paper cites Multi-step Jailbreaking Privacy Attacks on ChatGPT.

AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMs Multi-step Jailbreaking Privacy Attacks on ChatGPT

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-15T16:25:35.570589Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T16:25:35.570589Z digest=sha256:a4195627973b53d4ebe15a41503834eee3751d6ba4bd3b05a6cb02218515ce54

Observation c75d8f18-27c6-478c-aa7e-ace93bf76c7e · outbound

This paper cites The WMDP Benchmark: Measuring and Reducing Malicious Use With Unlearning.

AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMs The WMDP Benchmark: Measuring and Reducing Malicious Use With Unlearning

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-15T16:25:35.574328Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T16:25:35.574328Z digest=sha256:00f1a334c5b1b689f60325ceb5c4d2bdad6692f3ccd98be5138a9ca18e7656cd

Observation b91a5ea8-9ede-4d8a-87b6-4166cec6f27f · outbound

This paper cites TeleLoRA: Teleporting Model-Specific Alignment Across LLMs.

AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMs TeleLoRA: Teleporting Model-Specific Alignment Across LLMs

Reference 47

Resolution
verified exact
local_arxiv, observed 2026-08-15T16:25:36.529377Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-15T16:25:35.582429Z digest=sha256:ca5305d91533a7cbb44e94de9d2d982e0adcdb72d4fa2ff30d7562eeef363aff

Observation 2d16c294-5216-4998-9779-f7dbae85971a · outbound

This paper cites an unresolved cited work.

AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMs Unresolved cited work

Reference 48

Resolution
verified exact
raw_fallback, observed 2026-08-15T16:25:36.509939Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-15T16:25:35.586017Z digest=sha256:d387a6763b7b33b4c701a73e5c8fc6d529a599c9c9c6f57b09a8603f36f832d6

Observation 0e368ff1-ba39-4568-88d8-30414caf00aa · outbound

This paper cites DeepInception: Hypnotize Large Language Model to Be Jailbreaker.

AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMs DeepInception: Hypnotize Large Language Model to Be Jailbreaker

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-15T16:25:35.589876Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T16:25:35.589876Z digest=sha256:e9319d85bb546f0cb56e9b8793cd71f6dd87d8f88296c9acfb7e244cf671ef89

Observation c47c0b94-a64a-4edc-8d32-104d6ebbdb54 · outbound

This paper cites SHIELD: Evaluation and Defense Strategies for Copyright Compliance in LLM Text Generation.

AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMs SHIELD: Evaluation and Defense Strategies for Copyright Compliance in LLM Text Generation

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-15T16:25:35.597800Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T16:25:35.597800Z digest=sha256:7af8c43b0bc1fe0d7a65696c985051e64dedc707fcd8e852418150afac3a7297

Observation ca700848-885d-4c03-8a6f-32e5254d0726 · outbound

This paper cites AutoDAN: Generating Stealthy Jailbreak Prompts on Aligned Large Language Models.

AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMs AutoDAN: Generating Stealthy Jailbreak Prompts on Aligned Large Language Models

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-15T16:25:35.601622Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T16:25:35.601622Z digest=sha256:761b41024c0de97139493a238f6171f33fdba38a3ef687e44b40cc4f6f4ed405

Observation 98430e03-cf23-443b-9007-c53b22089ba1 · outbound

This paper cites HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust Refusal.

AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMs HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust Refusal

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-15T16:25:35.612838Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T16:25:35.612838Z digest=sha256:1494b3c6e7fbbf08084a88321348cfa4727a16eb7d3b23f48949e2a9d78435e9

Observation ca845d39-ef48-4bbc-9588-f5a40d6a5c14 · outbound

This paper cites R.; and Papernot, N.

AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMs R.; and Papernot, N

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-15T16:25:35.616456Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T16:25:35.616456Z digest=sha256:deefd1989f2db7c9b73415aa88bb1cd7ab77c6095c814fe025355ef9f90ef6fe

Observation 833a1260-eedb-4242-bd8a-bccd6a290aab · outbound

This paper cites Topic-Based Watermarks for Large Language Models.

AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMs Topic-Based Watermarks for Large Language Models

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-15T16:25:35.620427Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T16:25:35.620427Z digest=sha256:cfd33ff99ac3428aa5d2afaba0bf7469cdf0d394e88dbb3fa5719b5c71438da9

Observation eb29432f-9822-466a-b1fa-8fa53f9fb69f · outbound

This paper cites J.; Hassani, H.; Robey, A.; and Wong, E.

AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMs J.; Hassani, H.; Robey, A.; and Wong, E

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:25:37.165178Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-15T16:25:35.625006Z digest=sha256:96770602f087b1edffb732f0236016d67c399d6060bc36ebb20eaa88fd6ea537

Observation 0edacc3f-f76f-4b44-b7a7-038ce3b123f5 · outbound

This paper cites In-Context Unlearning: Language Models as Few Shot Unlearners.

AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMs In-Context Unlearning: Language Models as Few Shot Unlearners

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-15T16:25:35.629057Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T16:25:35.629057Z digest=sha256:cb5ffb675539669c3300814f77e2f14b4fbec618a1206c457528d4f290b09248

Observation c3ff32b1-5882-4710-a2ed-0495bdeabd70 · outbound

This paper cites an unresolved cited work.

AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMs Unresolved cited work

Reference 60

Resolution
unresolved
raw_fallback, observed 2026-08-15T16:25:37.148810Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-15T16:25:35.633450Z digest=sha256:840e854f0ee5913c3dc3aea2247e7d4ec53f6101b0ff16559741d5bb32463bbd

Observation ddceaa23-7b54-4660-a7df-0af32dfa13e8 · outbound

This paper cites LLM Defenses Are Not Robust to Multi-Turn Human Jailbreaks Yet.

AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMs LLM Defenses Are Not Robust to Multi-Turn Human Jailbreaks Yet

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-15T16:25:35.637678Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T16:25:35.637678Z digest=sha256:e0a2b5c5d92e3339262fc98064ba7e487dfcb1d896cac8d80df3ce6d6cb366fa

Observation 53b212db-a8ec-46fa-84a7-f7fea271f2ea · outbound

This paper cites Direct Preference Optimization: Your Language Model is Secretly a Reward Model.

AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMs Direct Preference Optimization: Your Language Model is Secretly a Reward Model

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-15T16:25:35.641624Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T16:25:35.641624Z digest=sha256:d8ab2d0f5a6605c163f234454ca977cf23f85e0ef74bf8dc682bae4a1cb0311a

Observation 3ce97506-2eff-4deb-b347-a0fcd6367b25 · outbound

This paper cites an unresolved cited work.

AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMs Unresolved cited work

Reference 63

Resolution
unresolved
raw_fallback, observed 2026-08-15T16:25:37.134159Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-15T16:25:35.645915Z digest=sha256:d43c513964a2c244cd1619b2478c8bf771415933cd1caeaf1f665209bdb89937

Observation e76398f4-5be7-4001-bc36-72747e6bb4ad · outbound

This paper cites XSTest: A Test Suite for Identifying Exaggerated Safety Behaviours in Large Language Models.

AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMs XSTest: A Test Suite for Identifying Exaggerated Safety Behaviours in Large Language Models

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-15T16:25:35.649342Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T16:25:35.649342Z digest=sha256:dc5861900c73284da30194343376858cc9507bb166fba727d2de55200549a5c4

Observation 0cf543b4-961c-4339-a167-43334dd940df · outbound

This paper cites an unresolved cited work.

AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMs Unresolved cited work

Reference 65

Resolution
unresolved
raw_fallback, observed 2026-08-15T16:25:37.122424Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-15T16:25:35.653241Z digest=sha256:9a1d9f0aaac5246eea00e67c4cb39271e31f55eb5d2c657014a3e8b26e5aadc5

Observation 529d0413-b7ce-4b21-969d-7682e60c144b · outbound

This paper cites Language Models Are Greedy Reasoners: A Systematic Formal Analysis of Chain-of-Thought.

AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMs Language Models Are Greedy Reasoners: A Systematic Formal Analysis of Chain-of-Thought

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-15T16:25:35.657150Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T16:25:35.657150Z digest=sha256:c9425aa9b245aa41f41a5ddf9bbdbbfed1a46e37f405719bb996bd7dd33e1fe1

Observation 0054103c-e053-4a9b-87c5-b95df03abe64 · outbound

This paper cites Scalable and Transferable Black-Box Jailbreaks for Language Models via Persona Modulation.

AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMs Scalable and Transferable Black-Box Jailbreaks for Language Models via Persona Modulation

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-15T16:25:35.661660Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T16:25:35.661660Z digest=sha256:3b99e0288e768018db76a4c1d277d99a45d404afe8ae34491253dfa0d0fd40b3

Observation 399fe1e5-aac0-4f27-a586-77d9448e3b04 · outbound

This paper cites an unresolved cited work.

AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMs Unresolved cited work

Reference 68

Resolution
unresolved
raw_fallback, observed 2026-08-15T16:25:37.110048Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-15T16:25:35.666208Z digest=sha256:3141b79f755e35aa5854b8c1b801a3f501e607196cb24231c9aece76163f25f3

Observation 36314a39-cde8-4e78-9130-60869cbd74c4 · outbound

This paper cites Survey of Vulnerabilities in Large Language Models Revealed by Adversarial Attacks.

AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMs Survey of Vulnerabilities in Large Language Models Revealed by Adversarial Attacks

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-15T16:25:35.669919Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T16:25:35.669919Z digest=sha256:b38c9c2ef9969a7e9793358d2f04869465fef53295d4a81f5bb3b9d1d1522f84

Observation 9ac3813f-9ba9-4c50-a3b5-403f03f45be1 · outbound

This paper cites Do Anything Now.

AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMs Do Anything Now

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:25:37.097746Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-15T16:25:35.673744Z digest=sha256:dae462ae0e68bd4f99f7fa3b9147ca5b57fa599a8e909dcd33e069efbe1b8cf4

Observation d1af0da0-2d5e-4a04-80e3-98009e17e73c · outbound

This paper cites "Do Anything Now": Characterizing and Evaluating In-The-Wild Jailbreak Prompts on Large Language Models.

AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMs "Do Anything Now": Characterizing and Evaluating In-The-Wild Jailbreak Prompts on Large Language Models

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-15T16:25:35.677147Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T16:25:35.677147Z digest=sha256:69ea852e2b2d3320a35d6ca1bea37a104a225a79bd36f1f1eeba83f158cb5d64

Observation 0df81bd0-f189-401c-b4da-e05e884a8bfa · outbound

This paper cites Counterfactual Explainable Incremental Prompt Attack Analysis on Large Language Models.

AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMs Counterfactual Explainable Incremental Prompt Attack Analysis on Large Language Models

Reference 72

Resolution
verified exact
local_arxiv, observed 2026-08-15T16:25:36.160505Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-15T16:25:35.681487Z digest=sha256:f1bfe9ab600af267dcdd56b89461511916f39db92fe4a4c268b32fbeffdb56e1

Observation 6d520293-7269-4cab-88c1-5d92ee86b72c · outbound

This paper cites A StrongREJECT for Empty Jailbreaks.

AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMs A StrongREJECT for Empty Jailbreaks

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-15T16:25:35.685837Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T16:25:35.685837Z digest=sha256:2e5b5f4bc38b358e12c0355dc1b4466c9dc926113a287bd8ee907246bf86978f

Observation b5dffa9a-35b4-4ccc-af4a-d3635da1ee57 · outbound

This paper cites Tamper-Resistant Safeguards for Open-Weight LLMs.

AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMs Tamper-Resistant Safeguards for Open-Weight LLMs

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-15T16:25:35.689623Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T16:25:35.689623Z digest=sha256:7f7b97927a5f1c4baad3aeef3bc11b9bcddf4882bff1f556030a291ca485ec96

Observation 8e6337e5-00e5-4a69-9bae-f25b6ef9dfa9 · outbound

This paper cites Gemma 3 Technical Report.

AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMs Gemma 3 Technical Report

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-15T16:25:35.694213Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T16:25:35.694213Z digest=sha256:62ae51c39885920aae764d6f74bcf10e084862a15efe295df26abd97b05dea6d

Observation 9af2336c-692a-4b0d-a3fa-f783b16507e6 · outbound

This paper cites an unresolved cited work.

AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMs Unresolved cited work

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-15T16:25:35.698672Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T16:25:35.698672Z digest=sha256:9b3c60ce8cc7b75b9b3d59e76b1f5afca83fabd4b471b5b568acd5deb6cc13d3

Observation 5655b3f3-d956-4b39-99a4-a1fa20e609ec · outbound

This paper cites Prompt, Divide, and Conquer: Bypassing Large Language Model Safety Filters via Segmented and Distributed Prompt Processing.

AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMs Prompt, Divide, and Conquer: Bypassing Large Language Model Safety Filters via Segmented and Distributed Prompt Processing

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-15T16:25:35.702135Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T16:25:35.702135Z digest=sha256:13716da2ce188205056b49456fd9c10f269f8a31a3a537d6ab9cf25ee0f931bf

Observation 874015b7-cb02-40d5-924a-cdcdf96b9abf · outbound

This paper cites Do-Not-Answer: A Dataset for Evaluating Safeguards in LLMs.

AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMs Do-Not-Answer: A Dataset for Evaluating Safeguards in LLMs

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-15T16:25:35.706234Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T16:25:35.706234Z digest=sha256:6d5d5a53eea0361571bc6aad5e7e76d75c0f40c1d62bd6c42510e2e5b2d212c6

Observation 51df8078-f20f-4b92-8d42-72156df20f57 · outbound

This paper cites Jailbroken: How Does LLM Safety Training Fail?.

AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMs Jailbroken: How Does LLM Safety Training Fail?

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-15T16:25:35.713916Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T16:25:35.713916Z digest=sha256:4edd69709462426df70ac0260abe8aa11c6f6e5e761fcd7f12dcd816348281ed

Observation 54f0fe5f-a4cd-4e81-974e-d7c8efcbbd7b · outbound

This paper cites SMILES-Prompting: A Novel Approach to LLM Jailbreak Attacks in Chemical Synthesis.

AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMs SMILES-Prompting: A Novel Approach to LLM Jailbreak Attacks in Chemical Synthesis

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-15T16:25:35.717645Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T16:25:35.717645Z digest=sha256:4682ccb43f46240c93a079b7e75ec2f512a6ba67720860ff7cedee06f24dafd6

Observation e8441d64-02c5-48ec-880f-9169d9c2e1e5 · outbound

This paper cites an unresolved cited work.

AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMs Unresolved cited work

Reference 82

Resolution
unresolved
raw_fallback, observed 2026-08-15T16:25:37.085522Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-15T16:25:35.721358Z digest=sha256:c93ce2284316e290443f40b6914e82d5cf83622316028f5d022831dc82c4e32d

Observation b33b4b48-4e00-4557-b1d1-80b2dd061767 · outbound

This paper cites an unresolved cited work.

AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMs Unresolved cited work

Reference 83

Resolution
unresolved
raw_fallback, observed 2026-08-15T16:25:37.069441Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-15T16:25:35.724987Z digest=sha256:b2061bd6fb05621b3ce9d6b6be23223ef3d875db6bc854c71bba6da5ca2836e8

Observation 7744cce7-13fd-4e4c-a439-05382bd6f77b · outbound

This paper cites RedAgent: Red Teaming Large Language Models with Context-aware Autonomous Language Agent.

AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMs RedAgent: Red Teaming Large Language Models with Context-aware Autonomous Language Agent

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-15T16:25:35.728642Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T16:25:35.728642Z digest=sha256:185601522d6db236cdaa1b4696b4bc665b8eec92cfee56325029294a1999d0de

Observation 6600dd76-3429-4677-9637-7463b1b07cf3 · outbound

This paper cites an unresolved cited work.

AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMs Unresolved cited work

Reference 85

Resolution
unresolved
raw_fallback, observed 2026-08-15T16:25:37.053075Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-15T16:25:35.732524Z digest=sha256:57f4a44b8e88fcdc1f44973e7a18a21494ad84a46065ebe29dedd129eec2da3b

Observation 602c5ab8-683e-42e2-b488-b0ba8192c37d · outbound

This paper cites Qwen3 Technical Report.

AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMs Qwen3 Technical Report

Reference 86

Resolution
unresolved
no resolver link, observed 2026-08-15T16:25:35.736230Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T16:25:35.736230Z digest=sha256:055951addb3bc9fa419943ea01b3d1e930fba22e779fd00f9865cdc3aaae6a5b

Observation 6737df63-0dd1-4039-b648-154aac2f6535 · outbound

This paper cites Jigsaw Puzzles: Splitting Harmful Questions to Jailbreak Large Language Models.

AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMs Jigsaw Puzzles: Splitting Harmful Questions to Jailbreak Large Language Models

Reference 87

Resolution
unresolved
no resolver link, observed 2026-08-15T16:25:35.740179Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T16:25:35.740179Z digest=sha256:6001e811c729704b7a61adc81578730782f7d3d645375e2a824597d43f9baad0

Observation 7bd9f140-c927-4bba-9625-210afbdebf6d · outbound

This paper cites The Dark Side of Trust: Authority Citation-Driven Jailbreak Attacks on Large Language Models.

AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMs The Dark Side of Trust: Authority Citation-Driven Jailbreak Attacks on Large Language Models

Reference 88

Resolution
unresolved
no resolver link, observed 2026-08-15T16:25:35.744416Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T16:25:35.744416Z digest=sha256:da6c2afadb82cd521f67e896df785a69362b28727748fcad8836b4aed348b96f

Observation a969486c-572c-44f7-aaab-60fbad43883b · outbound

This paper cites Low-Resource Languages Jailbreak GPT-4.

AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMs Low-Resource Languages Jailbreak GPT-4

Reference 89

Resolution
unresolved
no resolver link, observed 2026-08-15T16:25:35.748405Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T16:25:35.748405Z digest=sha256:a345981578c54392c2869dd30860ef3b9a2b463d92d605cb4f1de840559b95ba

Observation 6b76a273-6f9b-4777-8418-4253fb003343 · outbound

This paper cites Position: Editing Large Language Models Poses Serious Safety Risks.

AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMs Position: Editing Large Language Models Poses Serious Safety Risks

Reference 90

Resolution
unresolved
no resolver link, observed 2026-08-15T16:25:35.752237Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T16:25:35.752237Z digest=sha256:2569daa49613254e55526c6647f0afe6139982632f29d1ac688254a12cb70430

Observation c7c025be-ad8a-434b-ab31-d3acc0e091cb · outbound

This paper cites Don't Listen To Me: Understanding and Exploring Jailbreak Prompts of Large Language Models.

AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMs Don't Listen To Me: Understanding and Exploring Jailbreak Prompts of Large Language Models

Reference 92

Resolution
unresolved
no resolver link, observed 2026-08-15T16:25:35.760832Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T16:25:35.760832Z digest=sha256:39f1f178dfebeca327f9b9a19f208be6a4a428299e4f2e922fdb6bfeb3ea7548

Observation 2d1cdad1-23db-49bf-973b-2304ef7d8878 · outbound

This paper cites HellaSwag: Can a Machine Really Finish Your Sentence?.

AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMs HellaSwag: Can a Machine Really Finish Your Sentence?

Reference 93

Resolution
unresolved
no resolver link, observed 2026-08-15T16:25:35.765564Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T16:25:35.765564Z digest=sha256:b728599286b18e2a5b5adf71f000663bd502dd7b58ea7a4ab0ec58dbf8b3a06d

Observation f7125f4b-9524-4f2f-bb51-c7472aa3e70c · outbound

This paper cites How Johnny Can Persuade LLMs to Jailbreak Them: Rethinking Persuasion to Challenge AI Safety by Humanizing LLMs.

AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMs How Johnny Can Persuade LLMs to Jailbreak Them: Rethinking Persuasion to Challenge AI Safety by Humanizing LLMs

Reference 95

Resolution
unresolved
no resolver link, observed 2026-08-15T16:25:35.773425Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T16:25:35.773425Z digest=sha256:7aeea87b098dc8c57832c6ab390120d064fa0455e413876d516e9059f07bf0ea

Observation 51ede104-63ed-4a65-95df-a7068fb2edba · outbound

This paper cites LLM Embeddings Improve Test-time Adaptation to Tabular $Y|X$-Shifts.

AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMs LLM Embeddings Improve Test-time Adaptation to Tabular $Y|X$-Shifts

Reference 96

Resolution
verified exact
local_arxiv, observed 2026-08-15T16:25:35.868451Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-15T16:25:35.777182Z digest=sha256:41c5bc84491c5db12b98014f1237bb70ff54b21fc3bbb96ad350465e3e6e235a

Observation 7e84904e-ce8f-4cf2-b123-5d300a9c3e9d · outbound

This paper cites an unresolved cited work.

AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMs Unresolved cited work

Reference 97

Resolution
unresolved
raw_fallback, observed 2026-08-15T16:25:37.038890Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-15T16:25:35.781938Z digest=sha256:c9e726e2af7d53db4431b9925ac93c02402d072a15f33d4b7a60b2422ea8d1bc

Observation 7d03a866-8998-4115-942f-ca7cbea3c008 · outbound

This paper cites Making Them Ask and Answer: Jailbreaking Large Language Models in Few Queries via Disguise and Reconstruction.

AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMs Making Them Ask and Answer: Jailbreaking Large Language Models in Few Queries via Disguise and Reconstruction

Reference 98

Resolution
unresolved
no resolver link, observed 2026-08-15T16:25:35.786739Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T16:25:35.786739Z digest=sha256:352968f8b25f7498d6ec86787573ec3e970d16bfe28164b6d1594b47049a2d6e

Observation 2511efdb-5ff7-41fc-acc9-8bd3f5727caf · outbound

This paper cites LIMA: Less Is More for Alignment.

AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMs LIMA: Less Is More for Alignment

Reference 99

Resolution
unresolved
no resolver link, observed 2026-08-15T16:25:35.790232Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T16:25:35.790232Z digest=sha256:530c77c2d06ed3e1caf8324feb528ccbcaad27a640a15828af76dc90d0f2bbfb

Observation 82f303b1-65e0-4869-8233-fc05e3500b2e · outbound

This paper cites Aya Model: An Instruction Finetuned Open-Access Multilingual Language Model.

AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMs Aya Model: An Instruction Finetuned Open-Access Multilingual Language Model

Reference 100

Resolution
unresolved
no resolver link, observed 2026-08-15T16:25:35.794271Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T16:25:35.794271Z digest=sha256:0b59c05c0ee7f7054b41b0946a4d83744aa3b38d10f6055ae0355ca08348bcf9

Pith citing papers

Observation 9d9a0935-e55e-49d7-b954-564f6fade08c · inbound

Do Activation Monitors Survive Model Updates? Benchmarking, Predicting, and Repairing Activation-Monitor Staleness cites this paper.

Do Activation Monitors Survive Model Updates? Benchmarking, Predicting, and Repairing Activation-Monitor Staleness AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMs

Reference 54

Resolution
verified exact
arxiv_id, observed 2026-06-27T03:30:26.978682Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-06-27T03:24:24.714121Z digest=sha256:ec6d5579a1e441d92188bba516bd6480cc587fc0c882a7a14b5b3053a7ac4930