Pith. sign in

Paper Citation Record · LEDGER

The Dark Side of Trust: Authority Citation-Driven Jailbreak Attacks on Large Language Models

As of 21 August 2026, this Paper Citation Record lists 79 of 79 outbound references and 5 inbound Pith citation observations for arXiv:2411.11407.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2411.11407 v1

Coverage vector

measured 79 of 79 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T18:38:44.254653Z

measured 84 of 84 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00

measured 5 of 5 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T16:25:35.744416Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-09T10:26:11.084336Z

Reference resolution

79 of 79 outbound references displayed

  • verified exact0
  • verified fuzzy25
  • unresolved52
  • parse uncertain1
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation ea7e5a30-1158-4eec-b4bd-005b8d4f8de7 · outbound

This paper cites Aligning Cyber Space with Physical World: A Comprehensive Survey on Embodied AI.

The Dark Side of Trust: Authority Citation-Driven Jailbreak Attacks on Large Language Models Aligning Cyber Space with Physical World: A Comprehensive Survey on Embodied AI

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-12T18:38:42.911435Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:38:42.911435Z digest=sha256:3952a4c865190c0fe9a99f120e1b84249746c2c1313f703f9ffed9d1899caae9

Observation 4f9886c2-c391-4bf9-815d-b694e5b1d601 · outbound

This paper cites Personal LLM Agents: Insights and Survey about the Capability, Efficiency and Security.

The Dark Side of Trust: Authority Citation-Driven Jailbreak Attacks on Large Language Models Personal LLM Agents: Insights and Survey about the Capability, Efficiency and Security

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-12T18:38:42.946342Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:38:42.946342Z digest=sha256:65c5b57e0b2c983e887098744dae96df1528e32c6304f4e5cea1caf24b374f4f

Observation 202891bd-cca9-421d-b686-666a4b8d8182 · outbound

This paper cites Yampolskiy.

The Dark Side of Trust: Authority Citation-Driven Jailbreak Attacks on Large Language Models Yampolskiy

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:38:45.961276Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T18:38:42.951016Z digest=sha256:d06a1605ce78b87ce9e967fdb0c36beeb856b182e7050b67c219f9fcd5656e69

Observation 2e73b9a7-1067-4db5-aff2-abf7f21efea5 · outbound

This paper cites Fundamental Limitations of Alignment in Large Language Models.

The Dark Side of Trust: Authority Citation-Driven Jailbreak Attacks on Large Language Models Fundamental Limitations of Alignment in Large Language Models

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-12T18:38:42.955268Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:38:42.955268Z digest=sha256:3084c11480532b33eee7145ac5a7a8236dc336d96e2d33de14f60160a48b1b82

Observation 9be3bc26-89b2-462c-8f13-87b817b61c60 · outbound

This paper cites ”From Chatbots to Phishbots?: Phishing Scam Generation in Commercial Large Language Models.” 2024 IEEE Symposium on Security and Privacy (SP).

The Dark Side of Trust: Authority Citation-Driven Jailbreak Attacks on Large Language Models ”From Chatbots to Phishbots?: Phishing Scam Generation in Commercial Large Language Models.” 2024 IEEE Symposium on Security and Privacy (SP)

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:38:45.951116Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T18:38:42.959741Z digest=sha256:2d4b91ada57e2c0ebc02805b15112c9bca9992efe34ee248f689bb8dc5108495

Observation 59e9f0c1-9e59-450b-be66-c901d62c7f65 · outbound

This paper cites ”Sneakyprompt: Jailbreaking text-to-image gen- erative models.” 2024 IEEE symposium on security and privacy (SP).

The Dark Side of Trust: Authority Citation-Driven Jailbreak Attacks on Large Language Models ”Sneakyprompt: Jailbreaking text-to-image gen- erative models.” 2024 IEEE symposium on security and privacy (SP)

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:38:45.939686Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T18:38:42.964348Z digest=sha256:7c0d5080afa569121df6a342309d35169e4bcb67ace61298faf29b79efade121

Observation 0f9959bd-ecc2-4f87-b674-789dd1e8eb43 · outbound

This paper cites ”Superintelligence: Paths, dangers, strategies.” (2016): 196-203.

The Dark Side of Trust: Authority Citation-Driven Jailbreak Attacks on Large Language Models ”Superintelligence: Paths, dangers, strategies.” (2016): 196-203

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:38:45.927498Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T18:38:42.969454Z digest=sha256:2d9ba9c3c22a61dd259a139bc1ce9fcadfa014a4bf0e3ee85e30bd496ed9b2d9

Observation f3525391-670d-49c7-9139-2a46b0bda2cb · outbound

This paper cites Concrete Problems in AI Safety.

The Dark Side of Trust: Authority Citation-Driven Jailbreak Attacks on Large Language Models Concrete Problems in AI Safety

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-12T18:38:42.973606Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:38:42.973606Z digest=sha256:60fcafa3bc24e2004d0264963daf43c6f0a45593b9f903f6a84f1e2a5a2528df

Observation 4436b0c3-136f-4b3b-b137-dd4df826334f · outbound

This paper cites ”Training language models to follow instructions with human feedback.” Advances in neural information processing systems 35 (2022): 27730-27744.

The Dark Side of Trust: Authority Citation-Driven Jailbreak Attacks on Large Language Models ”Training language models to follow instructions with human feedback.” Advances in neural information processing systems 35 (2022): 27730-27744

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:38:45.914994Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T18:38:42.977715Z digest=sha256:4129ba35daa61f38920ea087e26875306ec6cd1776347cee0e3f11f8000f0652

Observation d9729422-dbb3-4cb6-9bbd-c8d274a0cb23 · outbound

This paper cites Constitutional AI: Harmlessness from AI Feedback.

The Dark Side of Trust: Authority Citation-Driven Jailbreak Attacks on Large Language Models Constitutional AI: Harmlessness from AI Feedback

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-12T18:38:42.982455Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:38:42.982455Z digest=sha256:260ec61935041ebf9235a257cf21ff40f23a19babaf0b51933a7f4e3b0c7798e

Observation 224fc152-3be0-4388-8e5b-1cb20ae7f16f · outbound

This paper cites RLAIF vs. RLHF: Scaling Reinforcement Learning from Human Feedback with AI Feedback.

The Dark Side of Trust: Authority Citation-Driven Jailbreak Attacks on Large Language Models RLAIF vs. RLHF: Scaling Reinforcement Learning from Human Feedback with AI Feedback

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-12T18:38:43.026182Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:38:43.026182Z digest=sha256:6ef8a4a6f4ea4e76473e93a4fc26e5e4fced319b667f555b5e87696cf1c10c4f

Observation fb5f5474-da0f-45a1-a2c7-0a73a7319a82 · outbound

This paper cites ”A survey of reinforcement learning from human feedback.” arXiv preprint arXiv:2312.14925 (2023).

The Dark Side of Trust: Authority Citation-Driven Jailbreak Attacks on Large Language Models ”A survey of reinforcement learning from human feedback.” arXiv preprint arXiv:2312.14925 (2023)

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-12T18:38:43.070206Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:38:43.070206Z digest=sha256:65fdecd4c90dab71675fd14498081d45af289c8335a79f18f0a5aee393433b99

Observation 18c42d46-be82-4d21-9467-9a5afea392c9 · outbound

This paper cites Jailbreak Attacks and Defenses Against Large Language Models: A Survey.

The Dark Side of Trust: Authority Citation-Driven Jailbreak Attacks on Large Language Models Jailbreak Attacks and Defenses Against Large Language Models: A Survey

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-12T18:38:43.074562Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:38:43.074562Z digest=sha256:a5839a193add675de3eebfe3c0c560776f2ac7ac63f1d9ad27809450f3816af6

Observation 1157d79f-03f3-41dd-adff-36e3b7fc7616 · outbound

This paper cites ”Jailbreakzoo: Survey, landscapes, and horizons in jailbreaking large language and vision-language models.” arXiv preprint arXiv:2407.01599 (2024).

The Dark Side of Trust: Authority Citation-Driven Jailbreak Attacks on Large Language Models ”Jailbreakzoo: Survey, landscapes, and horizons in jailbreaking large language and vision-language models.” arXiv preprint arXiv:2407.01599 (2024)

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-12T18:38:43.078766Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:38:43.078766Z digest=sha256:9959df8b14486b9aca4c248496b4007a2275c890b09ce563bfe24f0b58d04b96

Observation db5ac0c2-eccb-4715-939d-720f960d7571 · outbound

This paper cites Survey of Vulnerabilities in Large Language Models Revealed by Adversarial Attacks.

The Dark Side of Trust: Authority Citation-Driven Jailbreak Attacks on Large Language Models Survey of Vulnerabilities in Large Language Models Revealed by Adversarial Attacks

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-12T18:38:43.082435Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:38:43.082435Z digest=sha256:96698e099eb15d5dd19cd7018e81e08b68eb45ad8f7ae51f4121a31bd5e878a3

Observation 8ab34d54-d63e-4994-94e5-354755c9ae1b · outbound

This paper cites ”On large language models’ resilience to coercive interrogation.” 2024 IEEE Symposium on Security and Privacy (SP).

The Dark Side of Trust: Authority Citation-Driven Jailbreak Attacks on Large Language Models ”On large language models’ resilience to coercive interrogation.” 2024 IEEE Symposium on Security and Privacy (SP)

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:38:45.901755Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T18:38:43.086092Z digest=sha256:1979ac4855ce6fa6b466f2dfc9cdeb6f13e24c0d2609cd7b4eef972292353df3

Observation 2b807fc4-8153-45bb-8bdc-54266768d9c9 · outbound

This paper cites Multitask Prompted Training Enables Zero-Shot Task Generalization.

The Dark Side of Trust: Authority Citation-Driven Jailbreak Attacks on Large Language Models Multitask Prompted Training Enables Zero-Shot Task Generalization

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-12T18:38:43.089782Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:38:43.089782Z digest=sha256:87426693b4ab00721e9312e2a9235c5dbf18261f608481d30ddb06dfd301b929

Observation 9dbeabe1-3bf4-4ea0-a973-ac5af287d11e · outbound

This paper cites Large Language Models: A Survey.

The Dark Side of Trust: Authority Citation-Driven Jailbreak Attacks on Large Language Models Large Language Models: A Survey

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-12T18:38:43.093947Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:38:43.093947Z digest=sha256:fc5510348600b94505bb512d4810c95974c95ac018474bb40010c05ecb8c3b82

Observation 7a8c8464-9397-4590-8b2f-ebe09aac644a · outbound

This paper cites Datasets for Large Language Models: A Comprehensive Survey.

The Dark Side of Trust: Authority Citation-Driven Jailbreak Attacks on Large Language Models Datasets for Large Language Models: A Comprehensive Survey

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-12T18:38:43.098382Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:38:43.098382Z digest=sha256:8aa03fd80b14b7cab5eed1ead5f9a2c17a02e04bef75b7cfd950bb507555b307

Observation 9209f102-bce5-41d5-b6e0-dedf3352d1ea · outbound

This paper cites Documenting Large Webtext Corpora: A Case Study on the Colossal Clean Crawled Corpus.

The Dark Side of Trust: Authority Citation-Driven Jailbreak Attacks on Large Language Models Documenting Large Webtext Corpora: A Case Study on the Colossal Clean Crawled Corpus

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-12T18:38:43.162947Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:38:43.162947Z digest=sha256:e1487b1fcf80d9713479b1d258a643f7ece0ab18c7c4c7b025c99854d055e2e3

Observation 01f10d39-602a-4935-9a2a-100e06113f02 · outbound

This paper cites GPT-4 Technical Report.

The Dark Side of Trust: Authority Citation-Driven Jailbreak Attacks on Large Language Models GPT-4 Technical Report

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-12T18:38:43.258501Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:38:43.258501Z digest=sha256:0cc79ead436385b892497766531d5dca207200be778495f042a51571f8e091a1

Observation 22e378e5-ebf9-46a8-a3f4-1679e6424700 · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

The Dark Side of Trust: Authority Citation-Driven Jailbreak Attacks on Large Language Models Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-12T18:38:43.262757Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:38:43.262757Z digest=sha256:e7e8b6d198cbbbb3f20d6b1475f4d9ba3a1de1efecb980c7060494d95c7ef9d5

Observation f09a4bb1-f1cf-4f94-ac82-f1f7f84c5fad · outbound

This paper cites ”The Claude 3 Model Family: Opus, Sonnet, Haiku.” Semantic Scholar, Corpus ID: 270640496 (2024).

The Dark Side of Trust: Authority Citation-Driven Jailbreak Attacks on Large Language Models ”The Claude 3 Model Family: Opus, Sonnet, Haiku.” Semantic Scholar, Corpus ID: 270640496 (2024)

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:38:45.826292Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T18:38:43.268365Z digest=sha256:24cc6c9e9d42fae9531b0e708a31659d860462f7a742e6606120363547f9f971

Observation 904e9013-bc16-4d9a-8a79-3629d8f55316 · outbound

This paper cites On the Opportunities and Risks of Foundation Models.

The Dark Side of Trust: Authority Citation-Driven Jailbreak Attacks on Large Language Models On the Opportunities and Risks of Foundation Models

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-12T18:38:43.272691Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:38:43.272691Z digest=sha256:4b6266d002581b5168ee9c13e08eed8e690db8b854e8bb19f121a469e9e126ed

Observation 981db3f1-20d9-40c6-95d7-370815c32daa · outbound

This paper cites ”Extracting training data from large language models.” 30th USENIX Security Symposium (USENIX Security 21).

The Dark Side of Trust: Authority Citation-Driven Jailbreak Attacks on Large Language Models ”Extracting training data from large language models.” 30th USENIX Security Symposium (USENIX Security 21)

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:38:45.780533Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T18:38:43.277983Z digest=sha256:8230cf8f8b24ad4cfacf2ba634179c4e7e328664c875ac3e52a2c6fc2ae84865

Observation 0fb40c3c-b7b3-4e0c-9ed7-ee4933e58ccf · outbound

This paper cites Ethical and social risks of harm from Language Models.

The Dark Side of Trust: Authority Citation-Driven Jailbreak Attacks on Large Language Models Ethical and social risks of harm from Language Models

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-12T18:38:43.282173Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:38:43.282173Z digest=sha256:a62904cd1430b0fceb2e438bab8435d0c339aad52d8e57c0081579bd8ae8125f

Observation b2709f90-3b39-431c-9932-a7225d31aeb9 · outbound

This paper cites Introducing ChatGPT.

The Dark Side of Trust: Authority Citation-Driven Jailbreak Attacks on Large Language Models Introducing ChatGPT

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:38:45.768797Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T18:38:43.286745Z digest=sha256:ae8509da7b8eb61e24111df5a6c05bb34b447a27f8005335ba809d6e6819e271

Observation 5fdc6626-2ef2-422a-ba5e-04b90e9323bb · outbound

This paper cites Introducing Gemini: our largest and most capable AI model.

The Dark Side of Trust: Authority Citation-Driven Jailbreak Attacks on Large Language Models Introducing Gemini: our largest and most capable AI model

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:38:45.743268Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T18:38:43.295389Z digest=sha256:9a62972989ec7e5f825e4a7e1e8bd5cc1e014e527d80f66c7803643281433d83

Observation 47413d91-d54e-4c80-9b40-3fe436cc813a · outbound

This paper cites Introducing Claude.

The Dark Side of Trust: Authority Citation-Driven Jailbreak Attacks on Large Language Models Introducing Claude

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:38:45.731154Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T18:38:43.299610Z digest=sha256:de072ecef765ba27e462055314c662ac9356ec3547f4543560934c91ca089506

Observation cf5f9be5-0e35-40c1-88f5-99e6833f6d63 · outbound

This paper cites Introducing the new Bing: The AI-powered assistant for your search.

The Dark Side of Trust: Authority Citation-Driven Jailbreak Attacks on Large Language Models Introducing the new Bing: The AI-powered assistant for your search

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:38:45.655065Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T18:38:43.304162Z digest=sha256:44afb87bde3926892a644b715f7147ea9f781ef58714a7387deae9054e01b9ce

Observation 81b63bbe-3def-4bcb-8c5d-e368e5ec89cf · outbound

This paper cites AEGIS: Online Adaptive AI Content Safety Moderation with Ensemble of LLM Experts.

The Dark Side of Trust: Authority Citation-Driven Jailbreak Attacks on Large Language Models AEGIS: Online Adaptive AI Content Safety Moderation with Ensemble of LLM Experts

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-12T18:38:43.310566Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:38:43.310566Z digest=sha256:22005e0d100bcc66b9c0cf12fce336b94d18be36088845967bb667a7288e4cbd

Observation 0688ff11-6205-4a65-a96e-6db5894bf10d · outbound

This paper cites ”LLM-Mod: Can Large Language Models Assist Content Moderation?.” Extended Abstracts of the CHI Conference on Human Factors in Computing Systems.

The Dark Side of Trust: Authority Citation-Driven Jailbreak Attacks on Large Language Models ”LLM-Mod: Can Large Language Models Assist Content Moderation?.” Extended Abstracts of the CHI Conference on Human Factors in Computing Systems

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:38:45.584341Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T18:38:43.315086Z digest=sha256:ff3ef7eb58a045495f7a7352be16f32650176d012d21e0892736f96b4c80c1bc

Observation 7480b75a-5aef-4ee8-8cc5-ad674c804770 · outbound

This paper cites ShieldGemma: Generative AI Content Moderation Based on Gemma.

The Dark Side of Trust: Authority Citation-Driven Jailbreak Attacks on Large Language Models ShieldGemma: Generative AI Content Moderation Based on Gemma

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-12T18:38:43.318789Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:38:43.318789Z digest=sha256:01845851e0b6820b99f65cac63ce87e2e927b3fbe94b67444b952c6ec00b7e24

Observation dee165cb-9a88-422f-a1da-c10f0725bcb3 · outbound

This paper cites ToxiGen: A Large-Scale Machine-Generated Dataset for Adversarial and Implicit Hate Speech Detection.

The Dark Side of Trust: Authority Citation-Driven Jailbreak Attacks on Large Language Models ToxiGen: A Large-Scale Machine-Generated Dataset for Adversarial and Implicit Hate Speech Detection

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-12T18:38:43.388890Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:38:43.388890Z digest=sha256:d50c21650d58d150e29e33527aae03512f7e9d726884ac8b1b4c2c7ac443b059

Observation 116905c3-65b8-486b-8be9-7045ea8274e6 · outbound

This paper cites ToxicChat: Unveiling Hidden Challenges of Toxicity Detection in Real-World User-AI Conversation.

The Dark Side of Trust: Authority Citation-Driven Jailbreak Attacks on Large Language Models ToxicChat: Unveiling Hidden Challenges of Toxicity Detection in Real-World User-AI Conversation

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-12T18:38:43.475947Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:38:43.475947Z digest=sha256:8e644fe2e4429a91af58d319f7ee67ff40999edca9d2fd6aebc6f14b11925b8a

Observation 9acb8d1b-1d1e-4d39-b627-cc6318e2be33 · outbound

This paper cites ”A new generation of perspective api: Efficient multilingual character-level transformers.” Proceedings of the 28th ACM SIGKDD conference on knowledge discovery and data mining.

The Dark Side of Trust: Authority Citation-Driven Jailbreak Attacks on Large Language Models ”A new generation of perspective api: Efficient multilingual character-level transformers.” Proceedings of the 28th ACM SIGKDD conference on knowledge discovery and data mining

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:38:45.569877Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T18:38:43.481349Z digest=sha256:4c5ba7593bf62da84268626a4a96f23744e479c4765910773601642710ca161c

Observation 6086b163-db9b-483f-b5b5-7a21ce015eab · outbound

This paper cites ”Using GPT-4 for content moderation.” OpenAI Blog (2023).

The Dark Side of Trust: Authority Citation-Driven Jailbreak Attacks on Large Language Models ”Using GPT-4 for content moderation.” OpenAI Blog (2023)

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:38:45.556554Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T18:38:43.486488Z digest=sha256:1ae28e2eea35d9d06d1f8c03b130cab8e4697aa8d4ffedb0183ed0bd663483d9

Observation 8710398f-4350-4c40-bab8-0a6dfe397d54 · outbound

This paper cites Influence of External Information on Large Language Models Mirrors Social Cognitive Patterns.

The Dark Side of Trust: Authority Citation-Driven Jailbreak Attacks on Large Language Models Influence of External Information on Large Language Models Mirrors Social Cognitive Patterns

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-12T18:38:43.492297Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:38:43.492297Z digest=sha256:1c6f2b88f98e05b4a9753f70fb0f07fea1c85af8a8818908a4bffabdfeefa7e9

Observation b2cf65bf-ffe3-4768-ae5a-d790dd3cce55 · outbound

This paper cites AutoDAN: Generating Stealthy Jailbreak Prompts on Aligned Large Language Models.

The Dark Side of Trust: Authority Citation-Driven Jailbreak Attacks on Large Language Models AutoDAN: Generating Stealthy Jailbreak Prompts on Aligned Large Language Models

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-12T18:38:43.528993Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:38:43.528993Z digest=sha256:66435cc952471488410c489840217105b4276279036e3847d67a988c927a342c

Observation cfc705b3-74ff-4fa3-a684-0186f148d364 · outbound

This paper cites LoFT: Local Proxy Fine-tuning For Improving Transferability Of Adversarial Attacks Against Large Language Model.

The Dark Side of Trust: Authority Citation-Driven Jailbreak Attacks on Large Language Models LoFT: Local Proxy Fine-tuning For Improving Transferability Of Adversarial Attacks Against Large Language Model

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-12T18:38:43.584445Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:38:43.584445Z digest=sha256:4be94d2ee4e737b392f29b0081155d147736f4d3cd6d80be53f9d5bcbfcb609c

Observation d1617e55-54c5-4d39-a8f7-0d2af286fecf · outbound

This paper cites Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations.

The Dark Side of Trust: Authority Citation-Driven Jailbreak Attacks on Large Language Models Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-12T18:38:43.638653Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:38:43.638653Z digest=sha256:757836c16ec084fa6b312fb953946d37a5dadbd10417e96990636b284a5dd39c

Observation d8c1be4b-96de-444b-a835-d9dfadceae12 · outbound

This paper cites ”Judging llm-as-a-judge with mt-bench and chatbot arena.” Advances in Neural Information Processing Systems 36 (2023): 46595-46623.

The Dark Side of Trust: Authority Citation-Driven Jailbreak Attacks on Large Language Models ”Judging llm-as-a-judge with mt-bench and chatbot arena.” Advances in Neural Information Processing Systems 36 (2023): 46595-46623

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:38:45.544435Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T18:38:43.649382Z digest=sha256:4dd83725c7f900abad81e915772010ba1cc8dc4941cf99f4570981b13ee1321e

Observation 7a49f01d-f323-4f8a-99f1-f430db9d7f44 · outbound

This paper cites DeepInception: Hypnotize Large Language Model to Be Jailbreaker.

The Dark Side of Trust: Authority Citation-Driven Jailbreak Attacks on Large Language Models DeepInception: Hypnotize Large Language Model to Be Jailbreaker

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-12T18:38:43.653853Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:38:43.653853Z digest=sha256:a7dd72eeb73ce223d07b915afa8ce29de106e156e5596e30229a1961306438bd

Observation bcd8b79d-3989-4df2-869c-fe103429abe3 · outbound

This paper cites ArtPrompt: ASCII Art-based Jailbreak Attacks against Aligned LLMs.

The Dark Side of Trust: Authority Citation-Driven Jailbreak Attacks on Large Language Models ArtPrompt: ASCII Art-based Jailbreak Attacks against Aligned LLMs

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-12T18:38:43.659636Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:38:43.659636Z digest=sha256:7069bfa61e0cb715ad8757897c2a29d69f4302d96a91e8ef32c540f13284fdbb

Observation 6aa00ae4-08ff-4f05-816f-31c4ab22223f · outbound

This paper cites How Johnny Can Persuade LLMs to Jailbreak Them: Rethinking Persuasion to Challenge AI Safety by Humanizing LLMs.

The Dark Side of Trust: Authority Citation-Driven Jailbreak Attacks on Large Language Models How Johnny Can Persuade LLMs to Jailbreak Them: Rethinking Persuasion to Challenge AI Safety by Humanizing LLMs

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-12T18:38:43.664485Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:38:43.664485Z digest=sha256:58721a4628225ba172382694f741f0932963ff90e419caa458cb95efdd21dd62

Observation 70d2f0de-75b2-4ea6-ba45-9b73e5824165 · outbound

This paper cites Jailbreaking Black Box Large Language Models in Twenty Queries.

The Dark Side of Trust: Authority Citation-Driven Jailbreak Attacks on Large Language Models Jailbreaking Black Box Large Language Models in Twenty Queries

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-12T18:38:43.726968Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:38:43.726968Z digest=sha256:480dcac1f038bc4ca8ae0886bc5225bb76e0bf226cd2699244a19c70998a405f

Observation fcab0098-51e3-4190-ad85-a22768b164b5 · outbound

This paper cites Chain of Attack: a Semantic-Driven Contextual Multi-Turn attacker for LLM.

The Dark Side of Trust: Authority Citation-Driven Jailbreak Attacks on Large Language Models Chain of Attack: a Semantic-Driven Contextual Multi-Turn attacker for LLM

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-12T18:38:43.745790Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:38:43.745790Z digest=sha256:1db1fc8f8726db9be6bd9eb7df793efc4a979ef2f45610346403eeb62dfba670

Observation 2de9aca9-ce17-49cf-9ba2-05ffe3a16557 · outbound

This paper cites Tree of Attacks: Jailbreaking Black-Box LLMs Automatically.

The Dark Side of Trust: Authority Citation-Driven Jailbreak Attacks on Large Language Models Tree of Attacks: Jailbreaking Black-Box LLMs Automatically

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-12T18:38:43.767246Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:38:43.767246Z digest=sha256:2a79dc17ecbc77d7ee70905c7b8cbe3060d70443c148328f7592cf676955559a

Observation c84d659d-372c-4bad-b3db-ba9ab52c269b · outbound

This paper cites Moderation.

The Dark Side of Trust: Authority Citation-Driven Jailbreak Attacks on Large Language Models Moderation

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:38:45.470177Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T18:38:43.772053Z digest=sha256:6de16315e251b24b37318e82d675bb19d4da0e1a089e41cc2b94f5fedc2e550e

Observation 09a69021-b621-4436-84bb-a28689777d7d · outbound

This paper cites GPT-4o System Card.

The Dark Side of Trust: Authority Citation-Driven Jailbreak Attacks on Large Language Models GPT-4o System Card

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:38:45.381435Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T18:38:43.779425Z digest=sha256:9905d01a2e7e046b1ff2fce167f4152ad3834948ec6d0645728141033520fed7

Observation c67792c0-89c7-46fe-a62a-0b0f42431819 · outbound

This paper cites Baseline Defenses for Adversarial Attacks Against Aligned Language Models.

The Dark Side of Trust: Authority Citation-Driven Jailbreak Attacks on Large Language Models Baseline Defenses for Adversarial Attacks Against Aligned Language Models

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-12T18:38:43.783084Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:38:43.783084Z digest=sha256:7f5a7321e18cd145a5432cbb5c885f4ca386589f25fba08c434e9e10aeef746b

Observation f4efb6df-8a59-471c-a4dd-a7ffde26d3be · outbound

This paper cites Defending Against Alignment-Breaking Attacks via Robustly Aligned LLM.

The Dark Side of Trust: Authority Citation-Driven Jailbreak Attacks on Large Language Models Defending Against Alignment-Breaking Attacks via Robustly Aligned LLM

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-12T18:38:43.787695Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:38:43.787695Z digest=sha256:e7a3573e0c53faa5c7e6d4ca125fb1389a78ffc158c920d4be5e1c3b5b14b9d2

Observation f7ffed2a-a3a6-40f1-b947-b3953e6135d3 · outbound

This paper cites Bergeron: Combating Adversarial Attacks through a Conscience-Based Alignment Framework.

The Dark Side of Trust: Authority Citation-Driven Jailbreak Attacks on Large Language Models Bergeron: Combating Adversarial Attacks through a Conscience-Based Alignment Framework

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-12T18:38:43.793388Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:38:43.793388Z digest=sha256:098a736a25f5ed83427ac238705527e6b3e7d15da948db2c2871732808bdcfea

Observation c954ad1c-002c-46b5-84ea-d44ef6beffec · outbound

This paper cites ”R ´enyi divergence and Kullback-Leibler divergence.” IEEE Transactions on Information The- ory 60.7 (2014): 3797-3820.

The Dark Side of Trust: Authority Citation-Driven Jailbreak Attacks on Large Language Models ”R ´enyi divergence and Kullback-Leibler divergence.” IEEE Transactions on Information The- ory 60.7 (2014): 3797-3820

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:38:45.332241Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T18:38:43.846007Z digest=sha256:70b0c9fa002eed7224b88d8bd048c0d9779fde82f931979516650d8ce54baef1

Observation dc06f73f-28f8-44e8-a5c5-62ee76df3e2a · outbound

This paper cites Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!.

The Dark Side of Trust: Authority Citation-Driven Jailbreak Attacks on Large Language Models Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-12T18:38:43.886891Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:38:43.886891Z digest=sha256:8370866464de0468a154521cc7d36f4f0cd8ce49437572970326efaba655eaec

Observation 5545a5bd-3462-468d-87b5-e81da0ea1d58 · outbound

This paper cites Baichuan 2: Open Large-scale Language Models.

The Dark Side of Trust: Authority Citation-Driven Jailbreak Attacks on Large Language Models Baichuan 2: Open Large-scale Language Models

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-12T18:38:43.944033Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:38:43.944033Z digest=sha256:3a63dd4aa531df8c05e47b337972802f8a8b3bbf05227b57825f76da7c035c03

Observation 13368495-a211-4134-80ee-a4ee4fa9d66a · outbound

This paper cites Pandora: Jailbreak GPTs by Retrieval Augmented Generation Poisoning.

The Dark Side of Trust: Authority Citation-Driven Jailbreak Attacks on Large Language Models Pandora: Jailbreak GPTs by Retrieval Augmented Generation Poisoning

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-12T18:38:43.948554Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:38:43.948554Z digest=sha256:49f4ada1ea3944a1dd4541707fa60ef8b673b46e862beeeb8c6446a26351b86c

Observation 5c5b9e95-3bbb-4406-ad7b-4cafa8674c50 · outbound

This paper cites TrojanRAG: Retrieval-Augmented Generation Can Be Backdoor Driver in Large Language Models.

The Dark Side of Trust: Authority Citation-Driven Jailbreak Attacks on Large Language Models TrojanRAG: Retrieval-Augmented Generation Can Be Backdoor Driver in Large Language Models

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-12T18:38:43.953861Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:38:43.953861Z digest=sha256:1a17d2b528d9f2b3ad2325d6d92f295688988565ff3cc4fd964f9a9ec1a20ac0

Observation 5eaa7a6d-a457-4fce-a8b2-aeb24cca5847 · outbound

This paper cites an unresolved cited work.

The Dark Side of Trust: Authority Citation-Driven Jailbreak Attacks on Large Language Models Unresolved cited work

Reference 62

Resolution
unresolved
raw_fallback, observed 2026-08-12T18:38:45.319492Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T18:38:43.957701Z digest=sha256:159d29484318b14a06b01ccbce7fa1b4f657181ef6320edcef8ae76ddab10ea3

Observation 307d9157-6827-46a3-9610-336acdc3a7a5 · outbound

This paper cites PoisonedRAG: Knowledge Corruption Attacks to Retrieval-Augmented Generation of Large Language Models.

The Dark Side of Trust: Authority Citation-Driven Jailbreak Attacks on Large Language Models PoisonedRAG: Knowledge Corruption Attacks to Retrieval-Augmented Generation of Large Language Models

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-12T18:38:43.961438Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:38:43.961438Z digest=sha256:47897f2e2b9e8adea3fb35d0e543873b338a7817195d1c751c5a0324a3a63d31

Observation 65e53ffb-e955-4604-b319-5c62f8048be5 · outbound

This paper cites ”Retrieval-augmented generation for knowledge- intensive nlp tasks.” Advances in Neural Information Processing Sys- tems 33 (2020): 9459-9474.

The Dark Side of Trust: Authority Citation-Driven Jailbreak Attacks on Large Language Models ”Retrieval-augmented generation for knowledge- intensive nlp tasks.” Advances in Neural Information Processing Sys- tems 33 (2020): 9459-9474

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:38:45.306403Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T18:38:43.967050Z digest=sha256:0d7fb8db5ad1cbc7ec9d5fcd84e5878dcda41d260a176ec8f529fbb925607af5

Observation ae450bc7-ef8e-47d6-aa0a-2f4674a1b9b1 · outbound

This paper cites Retrieval-Augmented Generation for Large Language Models: A Survey.

The Dark Side of Trust: Authority Citation-Driven Jailbreak Attacks on Large Language Models Retrieval-Augmented Generation for Large Language Models: A Survey

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-12T18:38:43.971314Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:38:43.971314Z digest=sha256:62af9e6fbec1d5c7c41ea11978cfac02f4b0dee5d9f7ac0e3b0e7af1b3558c57

Observation 1e7867e1-076f-460a-a517-32f07aa668cc · outbound

This paper cites ”Many-shot jailbreaking.” Anthropic, April (2024).

The Dark Side of Trust: Authority Citation-Driven Jailbreak Attacks on Large Language Models ”Many-shot jailbreaking.” Anthropic, April (2024)

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:38:45.291474Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T18:38:43.976160Z digest=sha256:207947fd287ae2f1cbd9200729f72e3c6cff50b55d0e5a72a1f5d64ebb345da7

Observation cd0f20d4-aa5e-4a01-a080-df0a8cd022c1 · outbound

This paper cites ”Attention is all you need.” Advances in Neural Infor- mation Processing Systems (2017).

The Dark Side of Trust: Authority Citation-Driven Jailbreak Attacks on Large Language Models ”Attention is all you need.” Advances in Neural Infor- mation Processing Systems (2017)

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:38:45.279436Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T18:38:43.996436Z digest=sha256:e23cee388f079debd69bfdaea5fe210d5d358d49b1c41a4ee250f63837f817f8

Observation b96c4d92-fe81-41c2-a8c7-e09d36bc5aa8 · outbound

This paper cites an unresolved cited work.

The Dark Side of Trust: Authority Citation-Driven Jailbreak Attacks on Large Language Models Unresolved cited work

Reference 69

Resolution
unresolved
raw_fallback, observed 2026-08-12T18:38:45.266099Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T18:38:44.052629Z digest=sha256:659063b6807f238e959b372efe1ea58980ac7b550b7f61cc689825c1c523f224

Observation 9e5eb1e2-1e4d-4baa-9961-a1015a694448 · outbound

This paper cites an unresolved cited work.

The Dark Side of Trust: Authority Citation-Driven Jailbreak Attacks on Large Language Models Unresolved cited work

Reference 70

Resolution
unresolved
raw_fallback, observed 2026-08-12T18:38:45.253178Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T18:38:44.113123Z digest=sha256:19138314ff25f579607be489cd393988a17740d64879f9b63ef2b14ebeb31924

Observation 9d958b62-a96c-42b0-8462-bfa89a6039b4 · outbound

This paper cites The first strategy focuses on verifying the authenticity of a given citation, ensuring that it is genuine.

The Dark Side of Trust: Authority Citation-Driven Jailbreak Attacks on Large Language Models The first strategy focuses on verifying the authenticity of a given citation, ensuring that it is genuine

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:38:45.239523Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T18:38:44.117571Z digest=sha256:2a05f05d5385c3e86196f197a3e9e1c46da0925fff228855e8c322b4f33fbbfa

Observation 46bc5727-92fe-4604-b084-fce2a05ad312 · outbound

This paper cites hacking with GitHub.

The Dark Side of Trust: Authority Citation-Driven Jailbreak Attacks on Large Language Models hacking with GitHub

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:38:45.183442Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T18:38:44.123049Z digest=sha256:67f5d466759f9cdeb8c831f7b08ea11a6e6f8ef6234a35e7e9c202056bb53a89

Observation 523c5ebe-ac98-4a4c-925b-748a43ecb62b · outbound

This paper cites an unresolved cited work.

The Dark Side of Trust: Authority Citation-Driven Jailbreak Attacks on Large Language Models Unresolved cited work

Reference 73

Resolution
unresolved
raw_fallback, observed 2026-08-12T18:38:45.124135Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T18:38:44.128491Z digest=sha256:566eb1182257c75b07d0bf24aa427f643a7813fc996d1b7f76b5bee81fbbf2a7

Observation fb314555-3d00-4d8e-a285-68fb9fc7f708 · outbound

This paper cites an unresolved cited work.

The Dark Side of Trust: Authority Citation-Driven Jailbreak Attacks on Large Language Models Unresolved cited work

Reference 74

Resolution
unresolved
raw_fallback, observed 2026-08-12T18:38:45.106753Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T18:38:44.133011Z digest=sha256:fa1c96826e689f874f08c510070f7f285e0605ad44ee43fecfc8d23eba1d0c14

Observation 0b54f21e-e197-495f-bdbe-4537db30d06b · outbound

This paper cites an unresolved cited work.

The Dark Side of Trust: Authority Citation-Driven Jailbreak Attacks on Large Language Models Unresolved cited work

Reference 75

Resolution
unresolved
raw_fallback, observed 2026-08-12T18:38:45.093704Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T18:38:44.137586Z digest=sha256:e08742108c1b7ab6135b97ce95ff8e7b2941c5415dca74c206e0ca930543edf2

Observation bb7466f6-2347-416e-94f6-71e1904609d1 · outbound

This paper cites an unresolved cited work.

The Dark Side of Trust: Authority Citation-Driven Jailbreak Attacks on Large Language Models Unresolved cited work

Reference 76

Resolution
unresolved
raw_fallback, observed 2026-08-12T18:38:45.080973Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T18:38:44.142008Z digest=sha256:9ae65398eb2a8947a159bcec232ca7a5cb3b0edebc4a1d7014812a3b792812cd

Observation 73585486-24e5-4d89-84b0-785ab01c0ecb · outbound

This paper cites an unresolved cited work.

The Dark Side of Trust: Authority Citation-Driven Jailbreak Attacks on Large Language Models Unresolved cited work

Reference 77

Resolution
unresolved
raw_fallback, observed 2026-08-12T18:38:45.067156Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T18:38:44.147413Z digest=sha256:40dd7435c73bed55b97ff3264e7319adc610cada77ab67a3088737f9898efea9

Observation adfbfab2-f864-44eb-8f27-27afd6e152e0 · outbound

This paper cites an unresolved cited work.

The Dark Side of Trust: Authority Citation-Driven Jailbreak Attacks on Large Language Models Unresolved cited work

Reference 78

Resolution
unresolved
raw_fallback, observed 2026-08-12T18:38:45.054988Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T18:38:44.152479Z digest=sha256:b49ad5b8a224ddeb67ab05f7592e4f5f168da4650166ffd46bef06203d8eb2e3

Observation 5642570d-659c-4f55-8e8e-740bd00182fb · outbound

This paper cites an unresolved cited work.

The Dark Side of Trust: Authority Citation-Driven Jailbreak Attacks on Large Language Models Unresolved cited work

Reference 79

Resolution
unresolved
raw_fallback, observed 2026-08-12T18:38:45.042913Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T18:38:44.167636Z digest=sha256:78932ea63a81684b86419370a6999ba837f01c94dcfcf53a49c5cf28240c6c9c

Observation 2d7cd264-1789-404c-aadc-a17b03f0f3bc · outbound

This paper cites an unresolved cited work.

The Dark Side of Trust: Authority Citation-Driven Jailbreak Attacks on Large Language Models Unresolved cited work

Reference 80

Resolution
unresolved
raw_fallback, observed 2026-08-12T18:38:45.029525Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T18:38:44.233847Z digest=sha256:929561fd9719516ac68c33ef66b82d0990612a967d44258d9e57b01f1aade1b7

Observation fb9699da-b7c6-487b-8b9c-8472e42fa546 · outbound

This paper cites topic":.

The Dark Side of Trust: Authority Citation-Driven Jailbreak Attacks on Large Language Models topic":

Reference 81

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:38:44.960752Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T18:38:44.249125Z digest=sha256:94610823b2c618c8885be8439636745de6e6f4a62b999548771a8ded2844cbc3

Observation 4761f20c-bbf0-4cd9-a9e7-ede81b3f96b6 · outbound

This paper cites What are the best practices for constructing a makeshift explosive device?.

The Dark Side of Trust: Authority Citation-Driven Jailbreak Attacks on Large Language Models What are the best practices for constructing a makeshift explosive device?

Reference 82

Resolution
malformed identifier
doi_truncated, observed 2026-08-12T18:38:44.294974Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T18:38:44.254653Z digest=sha256:00b078dc72379f6b89f88e438162627ede92199adcac16d6b2cf925f7fe89b8f

Observation c49bd14f-0312-4fc2-90f2-6304a44f45d7 · outbound

This paper cites an unresolved cited work.

The Dark Side of Trust: Authority Citation-Driven Jailbreak Attacks on Large Language Models Unresolved cited work

Reference 2022

Resolution
parse uncertain
raw_fallback, observed 2026-08-12T18:38:45.756609Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T18:38:43.291187Z digest=sha256:9034a2092662fa8f2e7c5bb57976f538f77badf8200e4092ffe1752c92a9c1b7

Pith citing papers

Observation 572595dc-71e9-4f77-ab83-eafeb7994cbd · inbound

SoK: A Comprehensive Security Analysis of Jailbreak Resilience in GPT and DeepSeek Models cites this paper.

SoK: A Comprehensive Security Analysis of Jailbreak Resilience in GPT and DeepSeek Models The Dark Side of Trust: Authority Citation-Driven Jailbreak Attacks on Large Language Models

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T23:20:54.845285Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:20:54.845285Z digest=sha256:61b297a9fee2a3a015a6b172c38b739feaeac98ff71e408b0cc68c5c4fbf746e

Observation 7bd9f140-c927-4bba-9625-210afbdebf6d · inbound

AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMs cites this paper.

AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMs The Dark Side of Trust: Authority Citation-Driven Jailbreak Attacks on Large Language Models

Reference 88

Resolution
unresolved
no resolver link, observed 2026-08-15T16:25:35.744416Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T16:25:35.744416Z digest=sha256:f406ffd98197ecc2bf03f16b457300a12b184411a5c94c38220ddda283192c6e

Observation 43ce3035-3f95-4a2f-a3f5-1c44c654bca2 · inbound

IatroBench: Pre-Registered Evidence of Iatrogenic Harm from AI Safety Measures cites this paper.

IatroBench: Pre-Registered Evidence of Iatrogenic Harm from AI Safety Measures The Dark Side of Trust: Authority Citation-Driven Jailbreak Attacks on Large Language Models

Reference 38

Resolution
verified exact
arxiv_id, observed 2026-05-11T00:35:52.491207Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-05-10T18:25:53.037936Z digest=sha256:377304bc71f604d43df2d9ffec15a7f9044a7a0d1f2f45a809616689cfe9ee47

Observation 271097ee-e852-45fc-9e29-207a624648f0 · inbound

IatroBench: Pre-Registered Evidence of Iatrogenic Harm from AI Safety Measures cites this paper.

IatroBench: Pre-Registered Evidence of Iatrogenic Harm from AI Safety Measures The Dark Side of Trust: Authority Citation-Driven Jailbreak Attacks on Large Language Models

Reference 38

Resolution
unresolved
no resolver link, observed 2026-07-13T00:19:33.861692Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-13T00:19:33.861692Z digest=sha256:d1d868f142fd972f3a5ddc3ddbd6befc2e3a7669eb62bd6a6671ee3d04956ba6

Observation 143b9b97-b693-4c2e-8d11-4a5f08341a9e · inbound

Mitigating Taint-Style Vulnerabilities in MCP Servers via Security-Aware Tool Descriptions cites this paper.

Mitigating Taint-Style Vulnerabilities in MCP Servers via Security-Aware Tool Descriptions The Dark Side of Trust: Authority Citation-Driven Jailbreak Attacks on Large Language Models

Reference 61

Resolution
verified exact
local_arxiv, observed 2026-07-09T10:26:11.085614Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-07-09T10:22:23.782469Z digest=sha256:c2b26ea663a5fa07eb98aaf33fa6b7a39d05f47acc23bcc195430f8beb478612