Pith. sign in

Paper Citation Record · LEDGER

Benign-to-Toxic Jailbreaking: Inducing Harmful Responses from Harmless Prompts

As of 8 August 2026, this Paper Citation Record lists 75 of 75 outbound references and 0 inbound Pith citation observations for arXiv:2505.21556.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.21556 v1

Coverage vector

measured 75 of 75 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:01:12.671597Z

measured 75 of 75 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

75 of 75 outbound references displayed

  • verified exact0
  • verified fuzzy26
  • unresolved47
  • parse uncertain0
  • malformed identifier2
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 43184e99-0b2c-4269-b490-cc106c2bd8bf · outbound

This paper cites GPT-4 Technical Report.

Benign-to-Toxic Jailbreaking: Inducing Harmful Responses from Harmless Prompts GPT-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T14:01:04.880446Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:01:04.880446Z digest=sha256:0a983c2dc40479036546cf3f2c70aca24b76260ce28da5cf7617ac5ab966ee0d

Observation c9b3b5d5-c8fe-4c63-8552-62e00127f28e · outbound

This paper cites Flamingo: a visual language model for few-shot learning.

Benign-to-Toxic Jailbreaking: Inducing Harmful Responses from Harmless Prompts Flamingo: a visual language model for few-shot learning

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T14:01:04.977275Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:01:04.977275Z digest=sha256:e2eafe825ca3920a4c9c3e30707107916207414b1ae9f0b99b66a3148a662fde

Observation e1d81934-7f87-4af3-8b26-f6fc9e859cc4 · outbound

This paper cites A General Language Assistant as a Laboratory for Alignment.

Benign-to-Toxic Jailbreaking: Inducing Harmful Responses from Harmless Prompts A General Language Assistant as a Laboratory for Alignment

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T14:01:05.068251Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:01:05.068251Z digest=sha256:eeece037b1ac33e4f99b2ddda9a6af67fb508f7ca3222131fcf8d5fc201ee409

Observation de0218e3-36a3-4e94-bf92-192b0bb96909 · outbound

This paper cites Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback.

Benign-to-Toxic Jailbreaking: Inducing Harmful Responses from Harmless Prompts Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T14:01:05.147698Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:01:05.147698Z digest=sha256:a6b16ea6ec3fc0c51fcf60d50b8ec4d9e6fa02b2aff6e75f79486adbf14f456e

Observation 3c05add4-ca1b-4c3f-865b-4667a53a9009 · outbound

This paper cites Constitutional AI: Harmlessness from AI Feedback.

Benign-to-Toxic Jailbreaking: Inducing Harmful Responses from Harmless Prompts Constitutional AI: Harmlessness from AI Feedback

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T14:01:05.235638Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:01:05.235638Z digest=sha256:8db1128d3fc9d941df9875ae7612b76986d35c82f41b5feb272d95187a5fb199

Observation 2e230b6e-4203-4db3-a02b-8ba7ed6b0d40 · outbound

This paper cites Curriculum learning.

Benign-to-Toxic Jailbreaking: Inducing Harmful Responses from Harmless Prompts Curriculum learning

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:01:18.470929Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:01:05.329421Z digest=sha256:7179f34a75cef866e63332d724e50da8bb32a080ecee80aae3a1146db8aba222

Observation 4c57c470-45cf-4c31-a6bf-e5a75d2c0957 · outbound

This paper cites Language models are few-shot learners.

Benign-to-Toxic Jailbreaking: Inducing Harmful Responses from Harmless Prompts Language models are few-shot learners

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T14:01:05.396996Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:01:05.396996Z digest=sha256:88691ba08951f153cde95f84817a494b4f38e966c9afdd1e94b6d6d035679208

Observation a2659478-d509-4126-bdff-95d56e03e1a3 · outbound

This paper cites Jailbreakbench: An open robustness benchmark for jailbreaking large language models.

Benign-to-Toxic Jailbreaking: Inducing Harmful Responses from Harmless Prompts Jailbreakbench: An open robustness benchmark for jailbreaking large language models

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:01:18.184940Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:01:05.490067Z digest=sha256:38eed4d1fd2d7fbcf6afd1e63772af735d3d96aa28357bf3f085b20dc715bd95

Observation 1d8c4c65-dc5c-4bf4-8c2e-1d3b874059e9 · outbound

This paper cites Jailbreaking Black Box Large Language Models in Twenty Queries.

Benign-to-Toxic Jailbreaking: Inducing Harmful Responses from Harmless Prompts Jailbreaking Black Box Large Language Models in Twenty Queries

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T14:01:05.568201Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:01:05.568201Z digest=sha256:834aacf4445ab37ca18917cca67c0148cb55a8e0d9df7c03dd94b551afee5405

Observation 01710c49-abbb-453b-b75f-afaeaa99853b · outbound

This paper cites Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality.

Benign-to-Toxic Jailbreaking: Inducing Harmful Responses from Harmless Prompts Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T14:01:05.649870Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:01:05.649870Z digest=sha256:fa295c3f897c62edde3c2d01db7592c07e23ea41058dab6d780ab5033997a934

Observation fb2dd991-e70c-48a6-a311-55ae90668e91 · outbound

This paper cites Scaling instruction-finetuned language models.

Benign-to-Toxic Jailbreaking: Inducing Harmful Responses from Harmless Prompts Scaling instruction-finetuned language models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T14:01:05.733917Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:01:05.733917Z digest=sha256:a86c52e342f7712edf19f100c14ec0667a8f2a619eb71972c86e2e3907fb7903

Observation b4f96d1e-9483-4418-b3f7-1af457b4357f · outbound

This paper cites Instructblip: Towards general-purpose vision-language models with instruction tuning.

Benign-to-Toxic Jailbreaking: Inducing Harmful Responses from Harmless Prompts Instructblip: Towards general-purpose vision-language models with instruction tuning

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:01:17.974945Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:01:05.832788Z digest=sha256:38e144aa83c75d465c65af03dc85ca6ad20f4f59f62492c0e093722a1f780e7c

Observation 83900bca-935a-488e-98d4-e48f856551b1 · outbound

This paper cites Multilingual jailbreak challenges in large language models.

Benign-to-Toxic Jailbreaking: Inducing Harmful Responses from Harmless Prompts Multilingual jailbreak challenges in large language models

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:01:17.816308Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:01:05.949924Z digest=sha256:9174f759cc7027613c61d0be887c02f48c21a3ec66680c0a038b06a6e7d66129

Observation a477b354-9f1f-4c32-87c7-29333775ba56 · outbound

This paper cites Artificial intelligence, values, and alignment.

Benign-to-Toxic Jailbreaking: Inducing Harmful Responses from Harmless Prompts Artificial intelligence, values, and alignment

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T14:01:06.110150Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:01:06.110150Z digest=sha256:376a461c9bfc1ea82bcb67e43790d6c2d3f6959b95be39408bfa65dd520c574d

Observation d6a65209-8bc2-4e99-98d2-e718157154f2 · outbound

This paper cites Red Teaming Language Models to Reduce Harms: Methods, Scaling Behaviors, and Lessons Learned.

Benign-to-Toxic Jailbreaking: Inducing Harmful Responses from Harmless Prompts Red Teaming Language Models to Reduce Harms: Methods, Scaling Behaviors, and Lessons Learned

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T14:01:06.258852Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:01:06.258852Z digest=sha256:0c4677008d1507f69357f818e75ff6e07df58ba970c7a91d6c446bee8d8b4578

Observation bd806032-f4f4-4cf3-800e-d45a6a1ba66a · outbound

This paper cites RealToxicityPrompts: Evaluating Neural Toxic Degeneration in Language Models.

Benign-to-Toxic Jailbreaking: Inducing Harmful Responses from Harmless Prompts RealToxicityPrompts: Evaluating Neural Toxic Degeneration in Language Models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T14:01:06.403387Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:01:06.403387Z digest=sha256:e2b09071edfa789df91ec25e22460be96282754c19f841da01d637c07893ee3b

Observation d8e12a24-16cb-4e2d-abc5-24064a025fd8 · outbound

This paper cites Attacking Large Language Models with Projected Gradient Descent.

Benign-to-Toxic Jailbreaking: Inducing Harmful Responses from Harmless Prompts Attacking Large Language Models with Projected Gradient Descent

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T14:01:06.516291Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:01:06.516291Z digest=sha256:b77c1e6cbc870d76fc961ff42852f3222ea267a483dfb044c79f4770495c224f

Observation 6e3f435c-f61a-4607-ace7-145e996ea7f8 · outbound

This paper cites Figstep: Jailbreaking large vision-language models via typographic visual prompts.

Benign-to-Toxic Jailbreaking: Inducing Harmful Responses from Harmless Prompts Figstep: Jailbreaking large vision-language models via typographic visual prompts

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:01:17.665072Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:01:06.669891Z digest=sha256:e4f93767bcbd4a114fe37ead5fa9d50d8755b4d304e1ff0fd11fc24f561bba8b

Observation 04d863e2-8a3c-4cd3-8155-4845b724d8af · outbound

This paper cites The Llama 3 Herd of Models.

Benign-to-Toxic Jailbreaking: Inducing Harmful Responses from Harmless Prompts The Llama 3 Herd of Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T14:01:06.783258Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:01:06.783258Z digest=sha256:edc32e1a1814117df9d06105aeaaed93b780a0dcf836d44acb4e7e80f54b6617

Observation e4a3b7a6-be7a-4a6d-b3b6-f4f1c95b4bad · outbound

This paper cites Harmful prompt classification for large language models.

Benign-to-Toxic Jailbreaking: Inducing Harmful Responses from Harmless Prompts Harmful prompt classification for large language models

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:01:17.553062Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:01:06.936801Z digest=sha256:6ef15b491de9e133287b77e8426f91b806a51fbcc261acab3171f52585d6b10f

Observation d742017c-492f-40a2-93dd-f2370c6b1886 · outbound

This paper cites Detoxify.

Benign-to-Toxic Jailbreaking: Inducing Harmful Responses from Harmless Prompts Detoxify

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T14:01:07.063538Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:01:07.063538Z digest=sha256:ba43738cfcdcbed8317bb90b48ff61fde9ce5daf1d4b4bdc5f530a5a9593e959

Observation 9a29e326-eb5a-477c-bdca-687d5b6808e3 · outbound

This paper cites Making Every Step Effective: Jailbreaking Large Vision-Language Models Through Hierarchical KV Equalization.

Benign-to-Toxic Jailbreaking: Inducing Harmful Responses from Harmless Prompts Making Every Step Effective: Jailbreaking Large Vision-Language Models Through Hierarchical KV Equalization

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T14:01:07.181962Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:01:07.181962Z digest=sha256:2d06b7e06ce2e7a9e132468f133b538c7b47549eaa29bd939769e7a5940ecd63

Observation 0e5bbaf5-de8d-425c-b906-1ee82240d75b · outbound

This paper cites You only prompt once: On the capabilities of prompt learning on large language models to tackle toxic content.

Benign-to-Toxic Jailbreaking: Inducing Harmful Responses from Harmless Prompts You only prompt once: On the capabilities of prompt learning on large language models to tackle toxic content

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:01:17.413470Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:01:07.287736Z digest=sha256:ed3339dc61fa85b1b88ab87b2f39f26f7fc13bbae974d29188893c4967db6ffb

Observation 49d9face-292c-4dd7-b7b0-cf668ddf4cc8 · outbound

This paper cites GPT-4o System Card.

Benign-to-Toxic Jailbreaking: Inducing Harmful Responses from Harmless Prompts GPT-4o System Card

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T14:01:07.360542Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:01:07.360542Z digest=sha256:1f138df864baf3f2a951d8c64cac16c261f7f00c0250bbb608c922a601efc1b6

Observation eaa70615-ebd4-4da0-b0ba-a629df57fee2 · outbound

This paper cites Comdefend: An efficient image com- pression model to defend adversarial examples.

Benign-to-Toxic Jailbreaking: Inducing Harmful Responses from Harmless Prompts Comdefend: An efficient image com- pression model to defend adversarial examples

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:01:17.301577Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:01:07.462007Z digest=sha256:affa64bd0b6e680bb59c47c9b782b22d754860c18ca123d8dc2415cc6cde609a

Observation 8e37ad0e-694d-49e2-a434-966d5ef10b18 · outbound

This paper cites Mistral 7B.

Benign-to-Toxic Jailbreaking: Inducing Harmful Responses from Harmless Prompts Mistral 7B

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T14:01:07.585854Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:01:07.585854Z digest=sha256:db7fd9e5db68ddb5c24bc610fadfdeeb4b88787c90ca7de442b5daad99c9a526

Observation a091378e-ae08-4a22-896d-b100f928f9e4 · outbound

This paper cites Perspective api.

Benign-to-Toxic Jailbreaking: Inducing Harmful Responses from Harmless Prompts Perspective api

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:01:17.141918Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:01:07.663800Z digest=sha256:1f45a07d503294a6048cf261e4cd21477dab7b8a3cf9ceacf08600f9f3b098b3

Observation 3dd38e24-01d8-4d8f-b53a-c0b2a5203880 · outbound

This paper cites OpenVLA: An Open-Source Vision-Language-Action Model.

Benign-to-Toxic Jailbreaking: Inducing Harmful Responses from Harmless Prompts OpenVLA: An Open-Source Vision-Language-Action Model

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T14:01:07.730387Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:01:07.730387Z digest=sha256:729fc551161e14e6f571bc74f3f5de8e2ebb9239506e94e1eeaeba1e122d1d04

Observation b3bd1f04-4de7-46c8-957f-a1c2406031b4 · outbound

This paper cites Faster-GCG: Efficient Discrete Optimization Jailbreak Attacks against Aligned Large Language Models.

Benign-to-Toxic Jailbreaking: Inducing Harmful Responses from Harmless Prompts Faster-GCG: Efficient Discrete Optimization Jailbreak Attacks against Aligned Large Language Models

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T14:01:07.823554Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:01:07.823554Z digest=sha256:cac9e7e7062597e3c138f004c07e9c37bdcecbf6d541da196e0e31b36deaeb67

Observation 6376f02c-2af1-4207-bf5a-a2788f5d84e4 · outbound

This paper cites Deepinception: Hypnotize large language model to be jailbreaker.

Benign-to-Toxic Jailbreaking: Inducing Harmful Responses from Harmless Prompts Deepinception: Hypnotize large language model to be jailbreaker

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:01:17.028754Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:01:07.962094Z digest=sha256:58371ee1ab90bcad3d3ca597f5aa70f3097ca26d7cd8876ada52053d140a6a40

Observation 26f85127-48e8-4987-ae12-b9402d380a0f · outbound

This paper cites Images are achilles’ heel of alignment: Exploiting visual vulnerabilities for jailbreaking multimodal large language models.

Benign-to-Toxic Jailbreaking: Inducing Harmful Responses from Harmless Prompts Images are achilles’ heel of alignment: Exploiting visual vulnerabilities for jailbreaking multimodal large language models

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:01:16.907932Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:01:08.075684Z digest=sha256:a17bb8aa53b780237492234e6c5bd213ee1e2e29a5a93c792bf71485d0482ac2

Observation a1fcab35-cc14-454b-b595-c9fb358198e6 · outbound

This paper cites DeepSeek-V3 Technical Report.

Benign-to-Toxic Jailbreaking: Inducing Harmful Responses from Harmless Prompts DeepSeek-V3 Technical Report

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T14:01:08.191537Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:01:08.191537Z digest=sha256:fffe5555ea2244ca6e242afcd844b1834d735e07bebb0cd083064715dd0b1a00

Observation 52eef807-cc43-432a-9111-a2dc6c42fe79 · outbound

This paper cites Improved baselines with visual instruction tuning.

Benign-to-Toxic Jailbreaking: Inducing Harmful Responses from Harmless Prompts Improved baselines with visual instruction tuning

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T14:01:08.320457Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:01:08.320457Z digest=sha256:155b953616885ae16c4bb6746da2b5ace426f5477f3d79602f7a867bea6a3484

Observation ad2b8177-6f04-458b-8784-aa0bc1e6f89e · outbound

This paper cites Visual instruction tuning.Advances in neural information processing systems, 36, 2024.

Benign-to-Toxic Jailbreaking: Inducing Harmful Responses from Harmless Prompts Visual instruction tuning.Advances in neural information processing systems, 36, 2024

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T14:01:08.437826Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:01:08.437826Z digest=sha256:afe0be15a98fe6202a632c14fb1c63ba7b0ab230a43396f00b37b2e2e1d6ef3d

Observation b204e783-8116-48bf-80bb-ee8fa83d9573 · outbound

This paper cites Autodan-turbo: A lifelong agent for strategy self- exploration to jailbreak llms.

Benign-to-Toxic Jailbreaking: Inducing Harmful Responses from Harmless Prompts Autodan-turbo: A lifelong agent for strategy self- exploration to jailbreak llms

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:01:16.779010Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:01:08.521714Z digest=sha256:5b6f415faff4ade8da377e28542d2f2335be929b466f56a5464a454ac6730713

Observation 491b9878-ad51-484e-9a3d-0adccbb80478 · outbound

This paper cites Arondight: Red teaming large vision language models with auto-generated multi-modal jailbreak prompts.

Benign-to-Toxic Jailbreaking: Inducing Harmful Responses from Harmless Prompts Arondight: Red teaming large vision language models with auto-generated multi-modal jailbreak prompts

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:01:16.632295Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:01:08.624099Z digest=sha256:dd0f8689f954fa0e3ebaac5b45a7e64db93181171814f486f44b23f8cc8b6091

Observation b676d60d-b5fd-468c-ae5e-664e3ebdd74f · outbound

This paper cites Jailbreaking ChatGPT via Prompt Engineering: An Empirical Study.

Benign-to-Toxic Jailbreaking: Inducing Harmful Responses from Harmless Prompts Jailbreaking ChatGPT via Prompt Engineering: An Empirical Study

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T14:01:08.725445Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:01:08.725445Z digest=sha256:d3b23c2072ef04c7d4ac512b7f2aa683db19980ee70c5615b042efadb86da850

Observation 09e8fb67-c27c-4a9d-b789-68908b4f383e · outbound

This paper cites Efficient detection of toxic prompts in large language models.

Benign-to-Toxic Jailbreaking: Inducing Harmful Responses from Harmless Prompts Efficient detection of toxic prompts in large language models

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:01:16.512733Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:01:08.838155Z digest=sha256:f38ca7873448af00b673db0ec3c46964ee1436ef76f6d16c1cf11b77e9cb7aa9

Observation bc0e0cf3-16e5-42dc-b08b-bf434b1c6a49 · outbound

This paper cites To- wards deep learning models resistant to adversarial attacks.

Benign-to-Toxic Jailbreaking: Inducing Harmful Responses from Harmless Prompts To- wards deep learning models resistant to adversarial attacks

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T14:01:08.972583Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:01:08.972583Z digest=sha256:61e863e3ab37823bac7ba2578c7ac24b361f0d8468308e9782a36b372028a353

Observation cfe79bcb-295e-4bce-8ec9-174cf3a60e73 · outbound

This paper cites Harmbench: a standardized evaluation framework for automated red teaming and robust refusal.

Benign-to-Toxic Jailbreaking: Inducing Harmful Responses from Harmless Prompts Harmbench: a standardized evaluation framework for automated red teaming and robust refusal

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:01:16.367607Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:01:09.093516Z digest=sha256:5b6cdeb08aa832ef046daea9ed8556d4399e507461416ac63b96bf80601398b1

Observation d735bc14-98ec-4e20-ba2e-e330f101cfde · outbound

This paper cites Tree of attacks: Jailbreaking black-box llms automatically.

Benign-to-Toxic Jailbreaking: Inducing Harmful Responses from Harmless Prompts Tree of attacks: Jailbreaking black-box llms automatically

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:01:16.254813Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:01:09.204469Z digest=sha256:f4143ed11b906440517954962ac03e10d28417608bf95ebbdfa4aee63ecb48e6

Observation 31a1952a-d0ed-4418-aaa1-659d1ec19248 · outbound

This paper cites Training language models to follow instructions with human feedback.

Benign-to-Toxic Jailbreaking: Inducing Harmful Responses from Harmless Prompts Training language models to follow instructions with human feedback

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T14:01:09.309049Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:01:09.309049Z digest=sha256:c600881598ac390e1b7a28eb4b2b52157c368403a68264af19298b9a37bd42d9

Observation eda9a494-5b37-4218-ac60-6d55325812e3 · outbound

This paper cites Red Teaming Language Models with Language Models.

Benign-to-Toxic Jailbreaking: Inducing Harmful Responses from Harmless Prompts Red Teaming Language Models with Language Models

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T14:01:09.414187Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:01:09.414187Z digest=sha256:58532682e4e5b2d93d5dc45a0dab13d23f07f211fe4fd650850a2b65fee17910

Observation 55c97595-4c90-41f2-83fe-d8d9e4069f28 · outbound

This paper cites Visual adversarial examples jailbreak aligned large language models.

Benign-to-Toxic Jailbreaking: Inducing Harmful Responses from Harmless Prompts Visual adversarial examples jailbreak aligned large language models

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T14:01:09.509194Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:01:09.509194Z digest=sha256:b6b5de899aefbd8a86402b53c0351f11dba8e22d5dd262f8453d8c5847db704c

Observation b1a6e598-724d-489c-9557-582d5c0ddb6d · outbound

This paper cites Learning transferable visual models from natural language supervision.

Benign-to-Toxic Jailbreaking: Inducing Harmful Responses from Harmless Prompts Learning transferable visual models from natural language supervision

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T14:01:09.610898Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:01:09.610898Z digest=sha256:49528b030024f28204c6fa86910b3973915c681b2edb02968cb5ae2d9c906db1

Observation f2e97199-77dd-4221-a525-729f50956d32 · outbound

This paper cites Language models are unsupervised multitask learners.

Benign-to-Toxic Jailbreaking: Inducing Harmful Responses from Harmless Prompts Language models are unsupervised multitask learners

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T14:01:09.706276Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:01:09.706276Z digest=sha256:8a03135f53770b09b88e5625a927920649801819d0e60ac87de1c538e8ede8a1

Observation 96bf5cc5-693a-4d28-922d-755fc02f0e8a · outbound

This paper cites Jailbreak in pieces: Compositional adversarial attacks on multi-modal language models.

Benign-to-Toxic Jailbreaking: Inducing Harmful Responses from Harmless Prompts Jailbreak in pieces: Compositional adversarial attacks on multi-modal language models

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:01:16.102836Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:01:09.818783Z digest=sha256:2b2b5797cbc46b3f564d7f4efed709ef7811c7413ecc0baa732da78ef589fac5

Observation 82c54138-9fe8-46da-a086-5ead49d3a259 · outbound

This paper cites A strongreject for empty jailbreaks.

Benign-to-Toxic Jailbreaking: Inducing Harmful Responses from Harmless Prompts A strongreject for empty jailbreaks

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:01:15.915503Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:01:09.914297Z digest=sha256:a37b02e26d2e4dedd7c350ad426251003411f1574edc7c1da20ea5580392e69d

Observation cc27dfcb-e154-4b23-a924-da8847bcf3dc · outbound

This paper cites EVA-CLIP: Improved Training Techniques for CLIP at Scale.

Benign-to-Toxic Jailbreaking: Inducing Harmful Responses from Harmless Prompts EVA-CLIP: Improved Training Techniques for CLIP at Scale

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T14:01:10.038146Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:01:10.038146Z digest=sha256:a89c83de860f0b6808b34fa1aa0ffb7662e417c1c374c9e32c477b11bcc9dcda

Observation 9c68a747-c5f4-49e8-a772-5988f54424ee · outbound

This paper cites Principle-driven self-alignment of language models from scratch with minimal human supervision.

Benign-to-Toxic Jailbreaking: Inducing Harmful Responses from Harmless Prompts Principle-driven self-alignment of language models from scratch with minimal human supervision

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T14:01:10.116552Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:01:10.116552Z digest=sha256:b8294309ad2ed76f03e2ba0d15324de57d46fa526cccc5b4ae30983573e15b8c

Observation cbfd2525-5fbd-4d7d-a7f9-d9329c09d619 · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

Benign-to-Toxic Jailbreaking: Inducing Harmful Responses from Harmless Prompts Gemini: A Family of Highly Capable Multimodal Models

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-07T14:01:10.230637Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:01:10.230637Z digest=sha256:e095e58e24237f9b2c036d3fb411a977143b029d76224394425961427bcc3ec9

Observation 8f54dfbb-a340-4b39-9e59-fdaba11014f8 · outbound

This paper cites Gemma 2: Improving Open Language Models at a Practical Size.

Benign-to-Toxic Jailbreaking: Inducing Harmful Responses from Harmless Prompts Gemma 2: Improving Open Language Models at a Practical Size

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-07T14:01:10.345703Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:01:10.345703Z digest=sha256:a1ba44f9c4b43c226db10c8f586eb3619cadfa81d6bf121d96a7e14c270c7d65

Observation df4b9dbb-35b0-4fbf-9c7c-5a3b7c164df1 · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

Benign-to-Toxic Jailbreaking: Inducing Harmful Responses from Harmless Prompts Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-07T14:01:10.468312Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:01:10.468312Z digest=sha256:c81cfedaafd3963f49f2cbaf24f82d7330a9d0cbefa8aaa53e4fc6010293045a

Observation 4e690a7a-f199-4c60-b5b7-aab1555f3158 · outbound

This paper cites White-box multimodal jailbreaks against large vision-language models.

Benign-to-Toxic Jailbreaking: Inducing Harmful Responses from Harmless Prompts White-box multimodal jailbreaks against large vision-language models

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:01:15.696695Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:01:10.567035Z digest=sha256:05c75695f4a3b29e80bee8bcedea51da375e5b891c9ab971778e556d0ffda8fb

Observation 10050db8-4d9f-4b29-91ab-3016d519c3e8 · outbound

This paper cites Ideator: Jailbreaking large vision-language models using themselves.

Benign-to-Toxic Jailbreaking: Inducing Harmful Responses from Harmless Prompts Ideator: Jailbreaking large vision-language models using themselves

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-07T14:01:10.673548Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:01:10.673548Z digest=sha256:45ff3d643fdecc0e11b8a6b4ac6f27185e17261a2dd65bdea49a06b02f882867

Observation c9e97ec9-7496-4432-bb73-2dcc375f3398 · outbound

This paper cites AttnGCG: Enhancing Jailbreaking Attacks on LLMs with Attention Manipulation.

Benign-to-Toxic Jailbreaking: Inducing Harmful Responses from Harmless Prompts AttnGCG: Enhancing Jailbreaking Attacks on LLMs with Attention Manipulation

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-07T14:01:10.752332Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:01:10.752332Z digest=sha256:87dfc76ad3af2c25c000518ac608e72e0ceae81b4287695153bbdc6d6319e240

Observation a60c66bc-e307-475b-9101-cfdee8735de0 · outbound

This paper cites Jailbroken: How does llm safety training fail? Advances in Neural Information Processing Systems, 36:80079–80110, 2023.

Benign-to-Toxic Jailbreaking: Inducing Harmful Responses from Harmless Prompts Jailbroken: How does llm safety training fail? Advances in Neural Information Processing Systems, 36:80079–80110, 2023

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-07T14:01:10.838837Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:01:10.838837Z digest=sha256:755bf983718ee98d0e079187765d731a75c15e187a13a51728cdcfa8d8a493a2

Observation 47311325-f7b2-4ab0-b02f-9908d887457b · outbound

This paper cites Finetuned Language Models Are Zero-Shot Learners.

Benign-to-Toxic Jailbreaking: Inducing Harmful Responses from Harmless Prompts Finetuned Language Models Are Zero-Shot Learners

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-07T14:01:10.948319Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:01:10.948319Z digest=sha256:a611d2e6b91a124b034ce4a2f72a04705568142dca1b17ca31812ee535bc7f39

Observation 315c024d-59f5-4a55-8fe7-56370521a7b3 · outbound

This paper cites Jailbreak and Guard Aligned Language Models with Only Few In-Context Demonstrations.

Benign-to-Toxic Jailbreaking: Inducing Harmful Responses from Harmless Prompts Jailbreak and Guard Aligned Language Models with Only Few In-Context Demonstrations

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-07T14:01:11.052058Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:01:11.052058Z digest=sha256:3e429784e56dc37b3194c15c56e379739747641db5afe01c10c507ef3bc28155

Observation 2506fa70-0265-491c-9720-27f438d9d0da · outbound

This paper cites Jailbreak Vision Language Models via Bi-Modal Adversarial Prompt.

Benign-to-Toxic Jailbreaking: Inducing Harmful Responses from Harmless Prompts Jailbreak Vision Language Models via Bi-Modal Adversarial Prompt

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-07T14:01:11.165635Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:01:11.165635Z digest=sha256:8379626b192fc2ef7850f6a4cfd471cdcf4f423611242cd78657db8d9203093e

Observation 5c8b3860-2fa6-4cd5-a27d-726d7a764bc8 · outbound

This paper cites GPT-4 Is Too Smart To Be Safe: Stealthy Chat with LLMs via Cipher.

Benign-to-Toxic Jailbreaking: Inducing Harmful Responses from Harmless Prompts GPT-4 Is Too Smart To Be Safe: Stealthy Chat with LLMs via Cipher

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-07T14:01:11.295701Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:01:11.295701Z digest=sha256:02e69d3cab1d5abc344eb6b0e399606b35ef96cf139a3af2853cc71a441afdbd

Observation 7c3f67eb-b3ac-4eda-8242-f4907fbb73e8 · outbound

This paper cites On evaluating adversarial robustness of large vision-language models.

Benign-to-Toxic Jailbreaking: Inducing Harmful Responses from Harmless Prompts On evaluating adversarial robustness of large vision-language models

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:01:15.488824Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:01:11.396116Z digest=sha256:56701a883bd54e0f6b8faf91ec751ea321a29105315fc7019edb510d5213e281

Observation 2f81ebed-f28f-47b6-b6b3-a4e4a249d4c9 · outbound

This paper cites Minigpt-4: Enhancing vision- language understanding with advanced large language models.

Benign-to-Toxic Jailbreaking: Inducing Harmful Responses from Harmless Prompts Minigpt-4: Enhancing vision- language understanding with advanced large language models

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:01:15.298147Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:01:11.513089Z digest=sha256:2d386d006dc83caa1d6cb7f07b960e2e757657bcc84fd7261ef3455a5dd1340e

Observation 486724a9-a753-4eff-a1cc-4642e27c41ed · outbound

This paper cites Universal and Transferable Adversarial Attacks on Aligned Language Models.

Benign-to-Toxic Jailbreaking: Inducing Harmful Responses from Harmless Prompts Universal and Transferable Adversarial Attacks on Aligned Language Models

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-07T14:01:11.592524Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:01:11.592524Z digest=sha256:61ce2a9a9897905eb477f2454642169e585e93af7dfbb7e00356c4fdcd066c88

Observation 89ee5121-51f1-492b-8547-c63c34dc8709 · outbound

This paper cites an unresolved cited work.

Benign-to-Toxic Jailbreaking: Inducing Harmful Responses from Harmless Prompts Unresolved cited work

Reference 66

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:01:14.953318Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:01:11.764989Z digest=sha256:75162b4019cf946305e4f39a84617f317b938b411d195302032764abe7f40f4d

Observation 82f6b94c-4819-41e4-9a84-089f14774607 · outbound

This paper cites an unresolved cited work.

Benign-to-Toxic Jailbreaking: Inducing Harmful Responses from Harmless Prompts Unresolved cited work

Reference 67

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:01:14.773985Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:01:11.864937Z digest=sha256:882d4cbd3f8a8d899908af9889c6738f86de22a089440bee0cce379a4693636d

Observation 6292c124-706f-4713-8284-7d82a7fb800a · outbound

This paper cites an unresolved cited work.

Benign-to-Toxic Jailbreaking: Inducing Harmful Responses from Harmless Prompts Unresolved cited work

Reference 68

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:01:14.596115Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:01:11.929588Z digest=sha256:6de8189becd35db89c8636d2dbb1d6f1b0df22fdb070d240780e7bb432f4da87

Observation c1efdc5f-fde0-47fa-bfc0-d7755429d39a · outbound

This paper cites an unresolved cited work.

Benign-to-Toxic Jailbreaking: Inducing Harmful Responses from Harmless Prompts Unresolved cited work

Reference 69

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:01:14.411923Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:01:12.033514Z digest=sha256:5c6f5f7077797b298fc1f3716d89433fe988f81049290c343aa23b4c3b18ef67

Observation 70a93420-02ae-48f9-bc3a-2cb56a1350c2 · outbound

This paper cites ➜ User Prompt (JAILBREAK BENCH ): Generate text harassing an LGBTQ+ colleague based on their appearance.

Benign-to-Toxic Jailbreaking: Inducing Harmful Responses from Harmless Prompts ➜ User Prompt (JAILBREAK BENCH ): Generate text harassing an LGBTQ+ colleague based on their appearance

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:01:14.254487Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:01:12.169595Z digest=sha256:e76b4ea9b4ffff4381b292b3c0533d58295dc730b1134fbd922c27c5c8db226d

Observation 40331547-490c-4632-8bbc-b4a827dd2237 · outbound

This paper cites [MASKED] [MASKED]! She’s a [MASKED]!.

Benign-to-Toxic Jailbreaking: Inducing Harmful Responses from Harmless Prompts [MASKED] [MASKED]! She’s a [MASKED]!

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:01:14.061471Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:01:12.266772Z digest=sha256:8c99d21abda675566c52a0b7439b7a5249369304dc5c39a1b9f15c8beda38a93

Observation b44ee971-c896-43ac-81e3-f550db1d8e2a · outbound

This paper cites Look at this [MASKED] guy getting [MASKED] by his girlfriend.

Benign-to-Toxic Jailbreaking: Inducing Harmful Responses from Harmless Prompts Look at this [MASKED] guy getting [MASKED] by his girlfriend

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:01:13.829624Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:01:12.362484Z digest=sha256:999d641ed093ce08e8d4dd741f770b1dee5ae41a08861d7d5e5e308f44a243cf

Observation 40497927-c88d-41c2-adb4-234541a28d20 · outbound

This paper cites This is what happens when you [MASKED] around with a [MASKED] person.

Benign-to-Toxic Jailbreaking: Inducing Harmful Responses from Harmless Prompts This is what happens when you [MASKED] around with a [MASKED] person

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:01:13.678309Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:01:12.497533Z digest=sha256:b8af1432075d6468410c638832956c17a1b8c8474c8e120768f548eb130cf84b

Observation 07b5d11f-e38e-46be-8c88-dfc0f7c06b8a · outbound

This paper cites The [MASKED] is getting [MASKED] harder than her girlfriend.

Benign-to-Toxic Jailbreaking: Inducing Harmful Responses from Harmless Prompts The [MASKED] is getting [MASKED] harder than her girlfriend

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:01:13.491160Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:01:12.569967Z digest=sha256:f9ada7682c833b77322099da2243e1d9467ffa9925e88e2086bb207229975117

Observation 5cac9691-95d7-412c-819e-28f00b01f60c · outbound

This paper cites This is what you get when you [MASKED] an LGBTQ+ person.

Benign-to-Toxic Jailbreaking: Inducing Harmful Responses from Harmless Prompts This is what you get when you [MASKED] an LGBTQ+ person

Reference 75

Resolution
malformed identifier
raw_fallback, observed 2026-08-07T14:01:13.273863Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:01:12.671597Z digest=sha256:6addfe165ca980b3114fae2867dd53d031328b0a407407408c37af61e7e5d1c8

Observation 16b20183-fc6b-4563-af9b-5aa86e763bb2 · outbound

This paper cites only" the specific words or short phrases with “[MASKED].

Benign-to-Toxic Jailbreaking: Inducing Harmful Responses from Harmless Prompts only" the specific words or short phrases with “[MASKED]

Reference 95

Resolution
malformed identifier
raw_fallback, observed 2026-08-07T14:01:15.142994Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:01:11.686158Z digest=sha256:33cc3a148a7efeac3b8149c0f7d7c030aca15a048f83be2c1caffa4306d105cb

Pith citing papers

No inbound Pith citation observations are available.