Pith. sign in

Paper Citation Record · LEDGER

Refusal is Not Safety! Benchmarking Latent Safety Risks of LLM-Driven Content Humorization

As of 18 August 2026, this Paper Citation Record lists 38 of 38 outbound references and 0 inbound Pith citation observations for arXiv:2607.15977.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.15977 v1

Coverage vector

measured 38 of 38 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-01T21:47:02.992276Z

measured 38 of 38 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

38 of 38 outbound references displayed

  • verified exact8
  • verified fuzzy0
  • unresolved30
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation d00c10e2-a451-457f-a55b-b5162d1adc0b · outbound

This paper cites 2026 , isbn =.

Refusal is Not Safety! Benchmarking Latent Safety Risks of LLM-Driven Content Humorization 2026 , isbn =

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-01T21:46:59.214081Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T21:46:59.214081Z digest=sha256:9144cf1b4dc1827f08db48a7e711f6724c22eeea6867d4452a49b5ea29be4934

Observation da6bc5dc-c63b-40cb-b991-674561f72688 · outbound

This paper cites Proceedings of the AAAI Conference on Artificial Intelligence , volume=.

Refusal is Not Safety! Benchmarking Latent Safety Risks of LLM-Driven Content Humorization Proceedings of the AAAI Conference on Artificial Intelligence , volume=

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-01T21:46:59.250762Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T21:46:59.250762Z digest=sha256:049ce79df64b791bb5863d3efce833a2a92a9db9b19af7959dddc28273638ad4

Observation 8d0db242-5ba6-42ef-93f8-d354a5618946 · outbound

This paper cites and Jain, Lalit K and Nowak, Robert D and Mankoff, Bob and Zhang, Jifan.

Refusal is Not Safety! Benchmarking Latent Safety Risks of LLM-Driven Content Humorization and Jain, Lalit K and Nowak, Robert D and Mankoff, Bob and Zhang, Jifan

Reference 3

Resolution
verified exact
doi, observed 2026-08-01T21:48:23.550091Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-01T21:46:59.333964Z digest=sha256:8356ebea3859cb6ed45899874f960d9d65511234eec0d9b7bfc1dbd8491ff363

Observation 0d4e8f00-1de4-4374-99d9-950de50da173 · outbound

This paper cites Small But Funny: A Feedback-Driven Approach to Humor Distillation.

Refusal is Not Safety! Benchmarking Latent Safety Risks of LLM-Driven Content Humorization Small But Funny: A Feedback-Driven Approach to Humor Distillation

Reference 4

Resolution
verified exact
doi, observed 2026-08-01T21:48:23.247303Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-01T21:46:59.410963Z digest=sha256:6460a3bd349c077474aaf3aab47be946b7f2f17f549744723bc7db7b39d84979

Observation c251c4ea-5caf-4751-8fd6-9ccb20b9e448 · outbound

This paper cites Pun Unintended: LLM s and the Illusion of Humor Understanding.

Refusal is Not Safety! Benchmarking Latent Safety Risks of LLM-Driven Content Humorization Pun Unintended: LLM s and the Illusion of Humor Understanding

Reference 5

Resolution
verified exact
doi, observed 2026-08-01T21:48:22.969822Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-01T21:46:59.497085Z digest=sha256:42d3f19e933f33412b6fa41a2ec94b47f393235dab6cd689144612e33c9e0811

Observation e2b425c1-0c06-4467-a7bb-8fef9088760d · outbound

This paper cites 2024 , isbn =.

Refusal is Not Safety! Benchmarking Latent Safety Risks of LLM-Driven Content Humorization 2024 , isbn =

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-01T21:46:59.574639Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T21:46:59.574639Z digest=sha256:eb1069505bd041024cde328fe07caa4bc0ad563b17d21090cefb8a7490d5c27e

Observation 231cf2a2-8af6-485d-862b-e5f37f023f0e · outbound

This paper cites Engagement Undermines Safety: How Stereotypes and Toxicity Shape Humor in Language Models.

Refusal is Not Safety! Benchmarking Latent Safety Risks of LLM-Driven Content Humorization Engagement Undermines Safety: How Stereotypes and Toxicity Shape Humor in Language Models

Reference 7

Resolution
verified exact
doi, observed 2026-08-01T21:48:22.811132Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-01T21:46:59.659930Z digest=sha256:f8a9b0dabdf1c8c7d412d2952618c365e59b0e5cf57dc5404080b8a34649cc18

Observation 49451c19-92e7-4b23-b306-f3e8477008b5 · outbound

This paper cites arXiv preprint arXiv:2603.17759 , year=.

Refusal is Not Safety! Benchmarking Latent Safety Risks of LLM-Driven Content Humorization arXiv preprint arXiv:2603.17759 , year=

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-01T21:46:59.743705Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T21:46:59.743705Z digest=sha256:f9f85336505e4dcbe992e8e661c7dc17b493aed6f8b9d3a2cb2e3a47bf9901d2

Observation 8ad4d120-1488-4fef-81fc-319e04fd7fba · outbound

This paper cites The Thirteenth International Conference on Learning Representations , year=.

Refusal is Not Safety! Benchmarking Latent Safety Risks of LLM-Driven Content Humorization The Thirteenth International Conference on Learning Representations , year=

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-01T21:46:59.817854Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T21:46:59.817854Z digest=sha256:4ff15ff903b4e5cd043d4d34961cd40076414f43eebf5bc75809ac3ea6df06b3

Observation e2292e0e-af6e-4c70-8b09-25448e10a053 · outbound

This paper cites ICML 2026 , year=.

Refusal is Not Safety! Benchmarking Latent Safety Risks of LLM-Driven Content Humorization ICML 2026 , year=

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-01T21:46:59.902434Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T21:46:59.902434Z digest=sha256:532f084d522c3fe86f2ecf2e0c5240b2d81cd454940346456b0d8852f1686371

Observation 801be5ed-3fb4-4ca8-8b89-25ac015e97d1 · outbound

This paper cites Getting Serious about Humor: Crafting Humor Datasets with Unfunny Large Language Models.

Refusal is Not Safety! Benchmarking Latent Safety Risks of LLM-Driven Content Humorization Getting Serious about Humor: Crafting Humor Datasets with Unfunny Large Language Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-01T21:46:59.986828Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T21:46:59.986828Z digest=sha256:49e9b9d441d5f327a9eea20b9017004a2a5103bcb11b74db54667bcd33a9d9b8

Observation ac00da6f-e595-4d55-beb5-fc7b71980d0d · outbound

This paper cites Chumor 2.0: Towards Better Benchmarking C hinese Humor Understanding from (Ruo Zhi Ba).

Refusal is Not Safety! Benchmarking Latent Safety Risks of LLM-Driven Content Humorization Chumor 2.0: Towards Better Benchmarking C hinese Humor Understanding from (Ruo Zhi Ba)

Reference 12

Resolution
verified exact
doi, observed 2026-08-01T21:48:22.621662Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-01T21:47:00.071256Z digest=sha256:00dc68ad5cb2f4f8830abd3cd0855c0a92ac36518dd98ec362ab3bb1155857eb

Observation 2aae92ec-fea7-459d-b20e-aebebaf1df51 · outbound

This paper cites Chumor 1.0: A Truly Funny and Challenging Chinese Humor Understanding Dataset from Ruo Zhi Ba.

Refusal is Not Safety! Benchmarking Latent Safety Risks of LLM-Driven Content Humorization Chumor 1.0: A Truly Funny and Challenging Chinese Humor Understanding Dataset from Ruo Zhi Ba

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-01T21:47:00.171514Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T21:47:00.171514Z digest=sha256:c45ab5e9241ecfcdf11c85d5c18251eaae7d0b229449a7a552ad8bcd11286c9f

Observation 9ed88eb5-da00-4681-af60-152b06456e23 · outbound

This paper cites Proceedings of the AAAI Conference on Artificial Intelligence , volume=.

Refusal is Not Safety! Benchmarking Latent Safety Risks of LLM-Driven Content Humorization Proceedings of the AAAI Conference on Artificial Intelligence , volume=

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-01T21:47:00.251140Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T21:47:00.251140Z digest=sha256:8ca3f54f6199d4391521c61f354e7cd30abd3fdc7bd7d5528c3c0d1844b72dfc

Observation ed96807e-efa9-45ff-bbd0-8f69125ef509 · outbound

This paper cites Incomplete Prompt Jailbreaks in Large Language Models.

Refusal is Not Safety! Benchmarking Latent Safety Risks of LLM-Driven Content Humorization Incomplete Prompt Jailbreaks in Large Language Models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-01T21:47:00.391154Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T21:47:00.391154Z digest=sha256:70c4e2406eeb0902cbaeb37fcbe9666ddf2e6ad636064c17431494153dac9fa2

Observation a58fdeea-1000-457e-b6cc-23162859590d · outbound

This paper cites Adjacent Words, Divergent Intents: Jailbreaking Large Language Models via Task Concurrency , url =.

Refusal is Not Safety! Benchmarking Latent Safety Risks of LLM-Driven Content Humorization Adjacent Words, Divergent Intents: Jailbreaking Large Language Models via Task Concurrency , url =

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-01T21:47:00.562639Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T21:47:00.562639Z digest=sha256:23a6cffbba9272939743aa13ba1245c8a1dc7b28d708f594125f9e5cef07b56a

Observation 635f8b6e-abe6-4ef2-96b1-38a5f73bcde4 · outbound

This paper cites 2024 , isbn =.

Refusal is Not Safety! Benchmarking Latent Safety Risks of LLM-Driven Content Humorization 2024 , isbn =

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-01T21:47:00.681639Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T21:47:00.681639Z digest=sha256:4cdec9e5a729b2a4b4779f1383300d05718d95a379a1861f8064510320ea6bd9

Observation 4f7a9a9a-fd45-4bf4-9120-8fcf7dfd4105 · outbound

This paper cites 2025 , isbn =.

Refusal is Not Safety! Benchmarking Latent Safety Risks of LLM-Driven Content Humorization 2025 , isbn =

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-01T21:47:00.825299Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T21:47:00.825299Z digest=sha256:56e1bc889d816c8c94d68e15700a53159552143eac411c755c1c10b1d9c1fe2e

Observation 128e6b02-ed6d-4fb8-a677-bb1c3cea5503 · outbound

This paper cites 33rd USENIX Security Symposium (USENIX Security 24) , year =.

Refusal is Not Safety! Benchmarking Latent Safety Risks of LLM-Driven Content Humorization 33rd USENIX Security Symposium (USENIX Security 24) , year =

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-01T21:47:00.931734Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T21:47:00.931734Z digest=sha256:56da8de8169ff57b20ff3c8cc3722e92ee4f6280a2e96b89ec66701b5e408de8

Observation aaad7948-05c9-4d22-ae6c-be7e1c74d4af · outbound

This paper cites Rothblum , booktitle=.

Refusal is Not Safety! Benchmarking Latent Safety Risks of LLM-Driven Content Humorization Rothblum , booktitle=

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-01T21:47:01.045844Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T21:47:01.045844Z digest=sha256:8a4bd82360e761e5cb2eee52e1b06af86de5e5099cabe824eacaca76814c7f2d

Observation 3e05f6ce-6445-47fb-93ad-cdc56cadc595 · outbound

This paper cites an unresolved cited work.

Refusal is Not Safety! Benchmarking Latent Safety Risks of LLM-Driven Content Humorization Unresolved cited work

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-01T21:47:01.086898Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T21:47:01.086898Z digest=sha256:65e590cb22485139ea014a7702e74ff3a532b1641496b157bf59c948a578f25e

Observation 3be2e25c-c8fd-44d8-ac54-2c0cabb13eda · outbound

This paper cites gpt-oss-120b & gpt-oss-20b Model Card.

Refusal is Not Safety! Benchmarking Latent Safety Risks of LLM-Driven Content Humorization gpt-oss-120b & gpt-oss-20b Model Card

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-01T21:47:01.180010Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T21:47:01.180010Z digest=sha256:dfdd783183430b674be030f24d6a01d03af337ec8d49cc69a7b842b79513472a

Observation 62c8ff9b-b573-46e6-8cf0-087eeb2f3096 · outbound

This paper cites OS -Sentinel: Towards Safety-Enhanced Mobile GUI Agents via Hybrid Validation in Realistic Workflows.

Refusal is Not Safety! Benchmarking Latent Safety Risks of LLM-Driven Content Humorization OS -Sentinel: Towards Safety-Enhanced Mobile GUI Agents via Hybrid Validation in Realistic Workflows

Reference 23

Resolution
verified exact
doi, observed 2026-08-01T21:48:22.337470Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-01T21:47:01.262257Z digest=sha256:352784875a224ec632e4b69bda8e2ef4cad4e281c67bea2e524e96200bfabc4a

Observation a7f3c676-74ee-4de1-89bc-1db186f1c5d7 · outbound

This paper cites arXiv preprint arXiv:2601.10527 , year=.

Refusal is Not Safety! Benchmarking Latent Safety Risks of LLM-Driven Content Humorization arXiv preprint arXiv:2601.10527 , year=

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-01T21:47:01.375929Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T21:47:01.375929Z digest=sha256:78543c5d398a8dc4a00d971c0e2611ea5c2c615f0249bcdb7c06189d3c9d7edc

Observation 301d3e81-7332-4677-9117-8484f79131c8 · outbound

This paper cites Proceedings of the 34th USENIX Conference on Security Symposium , articleno =.

Refusal is Not Safety! Benchmarking Latent Safety Risks of LLM-Driven Content Humorization Proceedings of the 34th USENIX Conference on Security Symposium , articleno =

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-01T21:47:01.515989Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T21:47:01.515989Z digest=sha256:0438cc665d8ab897eefc475f1118226b6175a0084e2ace495b41694ec0e4a350

Observation f65f516e-f30c-45a8-9016-7cfe9257966c · outbound

This paper cites 34th USENIX Security Symposium (USENIX Security 25) , pages=.

Refusal is Not Safety! Benchmarking Latent Safety Risks of LLM-Driven Content Humorization 34th USENIX Security Symposium (USENIX Security 25) , pages=

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-01T21:47:01.644159Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T21:47:01.644159Z digest=sha256:772605ca69881e6118b0322d883d69da2f7bedef97d48c4ba0c84c7dda1b5a50

Observation 88c93aa0-0367-4240-95d8-dadbbe8514c0 · outbound

This paper cites TRACE: Task-Aware Adaptive Self-Evolving Agentic Jailbreaking.

Refusal is Not Safety! Benchmarking Latent Safety Risks of LLM-Driven Content Humorization TRACE: Task-Aware Adaptive Self-Evolving Agentic Jailbreaking

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-01T21:47:01.765950Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T21:47:01.765950Z digest=sha256:b5daf3f4594bf0eb4a29da53351b84c8e9ceb81ac4d7b962d5d4968c7bab2405

Observation 91f50f54-ab44-40d8-aa45-50da8411fe5b · outbound

This paper cites AgentHarm: A Benchmark for Measuring Harmfulness of.

Refusal is Not Safety! Benchmarking Latent Safety Risks of LLM-Driven Content Humorization AgentHarm: A Benchmark for Measuring Harmfulness of

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-01T21:47:01.924481Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T21:47:01.924481Z digest=sha256:3242985899a1a2350a3674f443118af2fb174a7c8a94d90c1933b0511e106ab0

Observation 9597f2b5-eaa0-4e58-8c84-75cc6ede34a5 · outbound

This paper cites Vulnerability of Large Language Models to Output Prefix Jailbreaks: Impact of Positions on Safety.

Refusal is Not Safety! Benchmarking Latent Safety Risks of LLM-Driven Content Humorization Vulnerability of Large Language Models to Output Prefix Jailbreaks: Impact of Positions on Safety

Reference 29

Resolution
verified exact
doi, observed 2026-08-01T21:48:22.101117Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-01T21:47:02.084533Z digest=sha256:a61aa609713e9a1db37b1c961b64ce1873ee78a369be23fa3ac6daeee4795960

Observation 640a0cd5-e7da-4f5e-86be-223175b8d369 · outbound

This paper cites Humans welcome to observe.

Refusal is Not Safety! Benchmarking Latent Safety Risks of LLM-Driven Content Humorization Humans welcome to observe

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-01T21:47:02.252833Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T21:47:02.252833Z digest=sha256:5dbb06a79f8b2a02e6e40f78a3231ade4ef68f5dbd747f6f1f2da41baba4d2ae

Observation 7e34d067-d2b7-4a47-9f7e-03fbbd7758c9 · outbound

This paper cites and Chu, Eric and Behbahani, Feryal and Faust, Aleksandra and Larochelle, Hugo , booktitle =.

Refusal is Not Safety! Benchmarking Latent Safety Risks of LLM-Driven Content Humorization and Chu, Eric and Behbahani, Feryal and Faust, Aleksandra and Larochelle, Hugo , booktitle =

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-01T21:47:02.391217Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T21:47:02.391217Z digest=sha256:fd157408cf10ed7887b27ca55c2ee5499d58a06e2dfe1075850aabd9c54c633e

Observation 47c1efde-1479-4c5f-a2c1-02e593962036 · outbound

This paper cites ``What do you call a dog that is incontrovertibly true? Dogma'': Testing LLM Generalization through Humor.

Refusal is Not Safety! Benchmarking Latent Safety Risks of LLM-Driven Content Humorization ``What do you call a dog that is incontrovertibly true? Dogma'': Testing LLM Generalization through Humor

Reference 32

Resolution
verified exact
doi, observed 2026-08-01T21:48:21.795206Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-01T21:47:02.405637Z digest=sha256:2acad8c4a7c44c48845e255b06bc81dc14111172f09f835907cd08fd30370160

Observation a62c85a0-2ba1-4ea2-9f02-c39b204b265c · outbound

This paper cites and Lee, Lillian and Da, Jeff and Zellers, Rowan and Mankoff, Robert and Choi, Yejin.

Refusal is Not Safety! Benchmarking Latent Safety Risks of LLM-Driven Content Humorization and Lee, Lillian and Da, Jeff and Zellers, Rowan and Mankoff, Robert and Choi, Yejin

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-01T21:47:02.433639Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T21:47:02.433639Z digest=sha256:2fcde904e760c1f9c65f2ab017b37eece2ae8ce693c52092a417da02903b16f2

Observation b53b31d0-2b3a-496a-9455-c5eea47bc95f · outbound

This paper cites PIG uard: Prompt Injection Guardrail via Mitigating Overdefense for Free.

Refusal is Not Safety! Benchmarking Latent Safety Risks of LLM-Driven Content Humorization PIG uard: Prompt Injection Guardrail via Mitigating Overdefense for Free

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-01T21:47:02.535422Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T21:47:02.535422Z digest=sha256:42a59e45b36cd64349f7ec3dc18747945b58ef29d911e8f6adc9c80401060ded

Observation 5097350a-49c0-4297-8af3-603e2da70b83 · outbound

This paper cites Justice or Prejudice? Quantifying Biases in LLM-as-a-Judge , url =.

Refusal is Not Safety! Benchmarking Latent Safety Risks of LLM-Driven Content Humorization Justice or Prejudice? Quantifying Biases in LLM-as-a-Judge , url =

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-01T21:47:02.618087Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T21:47:02.618087Z digest=sha256:d2654bd5283b9ff76df2d06221277fa06d2da58bca7fe9511bb0e543e74d8b0a

Observation 19955a6f-658c-4b33-a1af-67c02f4b52e5 · outbound

This paper cites Agent Data Injection Attacks are Realistic Threats to AI Agents.

Refusal is Not Safety! Benchmarking Latent Safety Risks of LLM-Driven Content Humorization Agent Data Injection Attacks are Realistic Threats to AI Agents

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-01T21:47:02.723794Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T21:47:02.723794Z digest=sha256:97e980bd10aea72765d16931f9feae82d425cbd4680a993632a61b984ad324f7

Observation cc7a7e8f-003e-4c28-90c8-d9ad024c0548 · outbound

This paper cites Proceedings of the 41st International Conference on Machine Learning , articleno =.

Refusal is Not Safety! Benchmarking Latent Safety Risks of LLM-Driven Content Humorization Proceedings of the 41st International Conference on Machine Learning , articleno =

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-01T21:47:02.881957Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T21:47:02.881957Z digest=sha256:3c785e5a8fd2279a78746611a026b9322e127309bf93c37a6356a1041765fc03

Observation 1bbdd3d9-0400-40a4-adf8-88faa7a8f1d2 · outbound

This paper cites Universal and Transferable Adversarial Attacks on Aligned Language Models.

Refusal is Not Safety! Benchmarking Latent Safety Risks of LLM-Driven Content Humorization Universal and Transferable Adversarial Attacks on Aligned Language Models

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-01T21:47:02.992276Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T21:47:02.992276Z digest=sha256:35db3158100ea943ee20e2eb27962d5e2eee6da740a0580cc6924f5e531b3722

Pith citing papers

No inbound Pith citation observations are available.