Pith. sign in

Paper Citation Record · LEDGER

Exploring the Vulnerability of the Content Moderation Guardrail in Large Language Models via Intent Manipulation

As of 8 August 2026, this Paper Citation Record lists 50 of 50 outbound references and 3 inbound Pith citation observations for arXiv:2505.18556.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.18556 v2

Coverage vector

measured 50 of 50 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:32:31.644019Z

measured 53 of 53 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-02T13:18:40.623601Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-18T17:51:41.885862Z

Reference resolution

50 of 50 outbound references displayed

  • verified exact1
  • verified fuzzy1
  • unresolved48
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation b431b791-0fd2-4ac2-90f8-61f550819cd4 · outbound

This paper cites URL: " 'urlintro :=.

Exploring the Vulnerability of the Content Moderation Guardrail in Large Language Models via Intent Manipulation URL: " 'urlintro :=

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T14:32:28.368220Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:32:28.368220Z digest=sha256:1dbcc6ff95732428c3a5e3205239a1669850258ecfa85bdd95bdf9e3e01a3d1e

Observation 3172e170-78cf-451b-a001-86059b6107a0 · outbound

This paper cites write newline.

Exploring the Vulnerability of the Content Moderation Guardrail in Large Language Models via Intent Manipulation write newline

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T14:32:28.460572Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:32:28.460572Z digest=sha256:ce4039cdfcce44eee8fb5a4d2faf6ffa0bf0a039ed46ac37d12d74df4d7025b8

Observation e41aa9c5-5712-43fd-88d4-784ad7b7964d · outbound

This paper cites an unresolved cited work.

Exploring the Vulnerability of the Content Moderation Guardrail in Large Language Models via Intent Manipulation Unresolved cited work

Reference 3

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:32:35.729169Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T14:32:28.546314Z digest=sha256:b225a6abe76d1f22b09979a990e83d6f9121de90ee9c4e52286ac7b3b4d8956f

Observation b126d2e1-42ec-4643-b165-ef19f7956778 · outbound

This paper cites an unresolved cited work.

Exploring the Vulnerability of the Content Moderation Guardrail in Large Language Models via Intent Manipulation Unresolved cited work

Reference 4

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:32:35.556435Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T14:32:28.583911Z digest=sha256:256184a9a2491ac9a80ea8f226da9607e2fc95ff2cf38c3a322100ff8e00473b

Observation 09c0e544-7a5d-481a-80b9-12e2c9a810fa · outbound

This paper cites an unresolved cited work.

Exploring the Vulnerability of the Content Moderation Guardrail in Large Language Models via Intent Manipulation Unresolved cited work

Reference 5

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:32:35.407729Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T14:32:28.655389Z digest=sha256:6c5028891a3e7bb3ad7e706f7ca65021c6af8cf9920f37128c20318858a54d1a

Observation 4ee2aaab-b997-4149-aaac-e5d97dde2032 · outbound

This paper cites Pappas, Florian Tramèr, Hamed Hassani, and Eric Wong.

Exploring the Vulnerability of the Content Moderation Guardrail in Large Language Models via Intent Manipulation Pappas, Florian Tramèr, Hamed Hassani, and Eric Wong

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T14:32:28.748572Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:32:28.748572Z digest=sha256:a525c861ab57b499389fc7fc1a4ac45fd8caff0c018d62806970f01a5c8ab3f2

Observation 506c0bed-abf0-423f-8d97-e31c7977d43e · outbound

This paper cites Jailbreaking Black Box Large Language Models in Twenty Queries.

Exploring the Vulnerability of the Content Moderation Guardrail in Large Language Models via Intent Manipulation Jailbreaking Black Box Large Language Models in Twenty Queries

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T14:32:28.798711Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:32:28.798711Z digest=sha256:7087afb727bc108248dce5ca2f0188b777d4850f525b88d9083fe6c95d768794

Observation 1b0ffcb1-bb92-4f76-a247-4ef24ff1250b · outbound

This paper cites Z-BERT-A: a zero-shot Pipeline for Unknown Intent detection.

Exploring the Vulnerability of the Content Moderation Guardrail in Large Language Models via Intent Manipulation Z-BERT-A: a zero-shot Pipeline for Unknown Intent detection

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-08-07T14:32:32.295390Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T14:32:28.849543Z digest=sha256:49238e072dc653bfc7e57b18ba5bef6513d7334d6b257f4eb921a84cdbca8bb9

Observation de1eedfa-e02d-4ed1-9394-bf001ae8372d · outbound

This paper cites an unresolved cited work.

Exploring the Vulnerability of the Content Moderation Guardrail in Large Language Models via Intent Manipulation Unresolved cited work

Reference 9

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:32:35.319299Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T14:32:28.949530Z digest=sha256:36bc09345ab89ec83433ac5f8805088799ff294f9f4cab0801ff7e2a591cbea0

Observation 0c92bd8b-b5ce-4772-aa71-25501a945267 · outbound

This paper cites an unresolved cited work.

Exploring the Vulnerability of the Content Moderation Guardrail in Large Language Models via Intent Manipulation Unresolved cited work

Reference 10

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:32:35.258130Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T14:32:29.030810Z digest=sha256:e0ee904ce47ff393ff139e8ea4f46acb7d9d8be6ff80a11c6981808ef0991163

Observation a2b84897-36d2-4d2b-84ef-9eaac98ee83e · outbound

This paper cites Deliberative Alignment: Reasoning Enables Safer Language Models.

Exploring the Vulnerability of the Content Moderation Guardrail in Large Language Models via Intent Manipulation Deliberative Alignment: Reasoning Enables Safer Language Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T14:32:29.074711Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:32:29.074711Z digest=sha256:24bf4e4753933e4e2d1fff5440d6416a19452b800695183be0519ff647415ac1

Observation 83ece4b4-ec12-41c9-9c8b-c471793193ec · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Exploring the Vulnerability of the Content Moderation Guardrail in Large Language Models via Intent Manipulation DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T14:32:29.133774Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:32:29.133774Z digest=sha256:c7af79708b4a7934b80240701ef849023819ca1b53d06118ca8dd11bc22ef222

Observation 31a650b5-5135-4391-bedc-1ad88c054921 · outbound

This paper cites an unresolved cited work.

Exploring the Vulnerability of the Content Moderation Guardrail in Large Language Models via Intent Manipulation Unresolved cited work

Reference 13

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:32:35.139141Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T14:32:29.224704Z digest=sha256:233d0b4eac85cc45ee20a0524c2447bdc959dcfd51f0bd6e7c47b82d325ae6a1

Observation 1a8a5e0b-ea4e-4b82-bd59-b5d088003154 · outbound

This paper cites GPT-4o System Card.

Exploring the Vulnerability of the Content Moderation Guardrail in Large Language Models via Intent Manipulation GPT-4o System Card

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T14:32:29.287134Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:32:29.287134Z digest=sha256:f674a76c0e98fb9d46923600ad846d8ecb19f33ed4eb256c292276a0bc00ebb0

Observation 882510a1-c0c5-42ec-88dd-476aca8704e5 · outbound

This paper cites OpenAI o1 System Card.

Exploring the Vulnerability of the Content Moderation Guardrail in Large Language Models via Intent Manipulation OpenAI o1 System Card

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T14:32:29.320086Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:32:29.320086Z digest=sha256:1841fae741643813586365028bfa2f0fd05c836782ca63a4202053caea12c714

Observation 5f64d5f7-277b-49ed-9b32-797732415312 · outbound

This paper cites an unresolved cited work.

Exploring the Vulnerability of the Content Moderation Guardrail in Large Language Models via Intent Manipulation Unresolved cited work

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T14:32:29.432174Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:32:29.432174Z digest=sha256:f7f1181881405c2a63f4a6f9070b0e7b31239cf4c53c2a2f768596908303e4f7

Observation df9b1b16-5935-4d78-86de-5a319a58074c · outbound

This paper cites an unresolved cited work.

Exploring the Vulnerability of the Content Moderation Guardrail in Large Language Models via Intent Manipulation Unresolved cited work

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T14:32:29.510906Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:32:29.510906Z digest=sha256:09ad535af4b42ae1e1bb731bc9f2a0a63f430ff1ae2c89595db39c3cf7bd495f

Observation 51e6f370-574e-401e-8292-f543d35becb0 · outbound

This paper cites an unresolved cited work.

Exploring the Vulnerability of the Content Moderation Guardrail in Large Language Models via Intent Manipulation Unresolved cited work

Reference 18

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:32:35.039835Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T14:32:29.539612Z digest=sha256:8005bd734476ac0cde7d396cc56c6d330e6ff345659c2072b6981efecb4481b5

Observation 81e35d6b-b051-47dc-b881-75a029c5c93b · outbound

This paper cites an unresolved cited work.

Exploring the Vulnerability of the Content Moderation Guardrail in Large Language Models via Intent Manipulation Unresolved cited work

Reference 19

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:32:34.913348Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T14:32:29.633976Z digest=sha256:01cd27716e839702eaaacbc642be13b94e3a1e0408a79af744392da741ccba4d

Observation 9820e81e-1a97-4abd-9d41-c424982aea51 · outbound

This paper cites Open Sesame! Universal Black Box Jailbreaking of Large Language Models.

Exploring the Vulnerability of the Content Moderation Guardrail in Large Language Models via Intent Manipulation Open Sesame! Universal Black Box Jailbreaking of Large Language Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T14:32:29.722321Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:32:29.722321Z digest=sha256:588320f5dcaeebcb924778461b1380e5979ab6101a0034871d20292a8215455d

Observation 3251d485-49a5-4b94-835f-62a66192a6df · outbound

This paper cites an unresolved cited work.

Exploring the Vulnerability of the Content Moderation Guardrail in Large Language Models via Intent Manipulation Unresolved cited work

Reference 21

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:32:34.776349Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T14:32:29.790056Z digest=sha256:e82032ea6b7fba5899226d2e6d352fac31d4126b18528ec292862d9b115aa83b

Observation 242c95c8-92d6-4104-b7ba-e7860536e439 · outbound

This paper cites DeepSeek-V3 Technical Report.

Exploring the Vulnerability of the Content Moderation Guardrail in Large Language Models via Intent Manipulation DeepSeek-V3 Technical Report

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T14:32:29.834692Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:32:29.834692Z digest=sha256:2581ff9897eb7b6bea0bb20f45862c2d4450c9c666693c1dcef0f2ec90db5929

Observation 87f4dd05-014a-4240-89c4-2321628b6b58 · outbound

This paper cites FlipAttack: Jailbreak LLMs via Flipping.

Exploring the Vulnerability of the Content Moderation Guardrail in Large Language Models via Intent Manipulation FlipAttack: Jailbreak LLMs via Flipping

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T14:32:29.915875Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:32:29.915875Z digest=sha256:50a610fafac28a1bfbdddfb1a3776e2ecf417b139065e26f6ee437fd24fd2419

Observation 0d61b577-7030-4b17-a34d-833315d29f99 · outbound

This paper cites an unresolved cited work.

Exploring the Vulnerability of the Content Moderation Guardrail in Large Language Models via Intent Manipulation Unresolved cited work

Reference 24

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:32:34.706179Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T14:32:29.970460Z digest=sha256:76d442ffeaab078826ff2423261c92cc187bcf5e8da42858d7f7b1bb9503507a

Observation 5991bb0f-c38b-479b-86b3-021d59dccfc8 · outbound

This paper cites HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust Refusal.

Exploring the Vulnerability of the Content Moderation Guardrail in Large Language Models via Intent Manipulation HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust Refusal

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T14:32:30.021994Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:32:30.021994Z digest=sha256:84eb2bbc988a91f0458eeacbb537fff86555ffcef616b8ffd89ae78396a40f61

Observation a8e1c95f-3cdb-4dd7-939e-16de1224f06f · outbound

This paper cites an unresolved cited work.

Exploring the Vulnerability of the Content Moderation Guardrail in Large Language Models via Intent Manipulation Unresolved cited work

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T14:32:30.064116Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:32:30.064116Z digest=sha256:001fe81305b6d275007b4dd598416b794904dbed3d17c5339abd50b1f282464c

Observation 4601b9e5-8120-418b-a1fd-0461eb6ece5c · outbound

This paper cites an unresolved cited work.

Exploring the Vulnerability of the Content Moderation Guardrail in Large Language Models via Intent Manipulation Unresolved cited work

Reference 27

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:32:34.577088Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T14:32:30.138925Z digest=sha256:69cfde3c620107f5d790159bc2ebbb0ff0a7178871672ba0b20615dd753a9738

Observation bd04e60b-495f-4d0b-859a-a2c0b99cf229 · outbound

This paper cites an unresolved cited work.

Exploring the Vulnerability of the Content Moderation Guardrail in Large Language Models via Intent Manipulation Unresolved cited work

Reference 28

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:32:34.464284Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T14:32:30.229495Z digest=sha256:665e7f9740759ed6b8f5f8e3dc407c90119edf4c9f9659db30631c9736b40886

Observation 48f16cf2-8790-4d24-b2d3-f2a00cc4603e · outbound

This paper cites an unresolved cited work.

Exploring the Vulnerability of the Content Moderation Guardrail in Large Language Models via Intent Manipulation Unresolved cited work

Reference 29

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:32:34.372707Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T14:32:30.267198Z digest=sha256:02c7167c561e4f220810cd8ebcc7b0d10d3d5096554373a1a3a890c9c6e9e8d0

Observation 7f88d5ee-0bcf-4695-9c52-275fba0d9acb · outbound

This paper cites IntentGPT: Few-shot Intent Discovery with Large Language Models.

Exploring the Vulnerability of the Content Moderation Guardrail in Large Language Models via Intent Manipulation IntentGPT: Few-shot Intent Discovery with Large Language Models

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T14:32:30.352898Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:32:30.352898Z digest=sha256:f6321382ad13ffb415a12f70988bb77237db7c7f214543c676e023e0bbbd750d

Observation 38ce1dea-3956-4aaf-bb0c-193073a598e0 · outbound

This paper cites an unresolved cited work.

Exploring the Vulnerability of the Content Moderation Guardrail in Large Language Models via Intent Manipulation Unresolved cited work

Reference 31

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:32:34.218954Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T14:32:30.432860Z digest=sha256:b6c8b0013eddeecf1893ea22420dfa12b041ea0e6f9167485baeceeb94b987d1

Observation 3d3b302f-04de-4aeb-ad6b-60458330f346 · outbound

This paper cites an unresolved cited work.

Exploring the Vulnerability of the Content Moderation Guardrail in Large Language Models via Intent Manipulation Unresolved cited work

Reference 32

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:32:34.141225Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T14:32:30.483183Z digest=sha256:018b757affec567117fd525e27e2f0fef2b39f5c4105a56419900675c0483802

Observation 003315bc-28a4-43f6-b6d9-af64111e6ab6 · outbound

This paper cites an unresolved cited work.

Exploring the Vulnerability of the Content Moderation Guardrail in Large Language Models via Intent Manipulation Unresolved cited work

Reference 33

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:32:34.031105Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T14:32:30.534650Z digest=sha256:cead06cb52aa71ef217f279c6f2b8edcc6a62472966f0d654fca78775b123fe9

Observation ef5bb73a-5f1c-4140-ab2f-96bfc7a71cc3 · outbound

This paper cites an unresolved cited work.

Exploring the Vulnerability of the Content Moderation Guardrail in Large Language Models via Intent Manipulation Unresolved cited work

Reference 34

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:32:33.904284Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T14:32:30.565324Z digest=sha256:4c11d048fdaaf4c92eef551601a5382a718c66253d01d2d3c1bac5f0dde07e3f

Observation 61305ab5-bfa8-431c-a5a8-6c4fadff7ba4 · outbound

This paper cites an unresolved cited work.

Exploring the Vulnerability of the Content Moderation Guardrail in Large Language Models via Intent Manipulation Unresolved cited work

Reference 35

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:32:33.820745Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T14:32:30.652851Z digest=sha256:6493cf40fceecae6cd09bd6a7197f3713088f6097371a036448150de40925896

Observation 50dc41cf-1536-4e7b-869a-64ad6d1d0615 · outbound

This paper cites Jailbreak and Guard Aligned Language Models with Only Few In-Context Demonstrations.

Exploring the Vulnerability of the Content Moderation Guardrail in Large Language Models via Intent Manipulation Jailbreak and Guard Aligned Language Models with Only Few In-Context Demonstrations

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T14:32:30.720380Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:32:30.720380Z digest=sha256:8c3d62be5e6eef03026b2755ca2ab2bd4edca63fb7e97b6f8e7894d8e825cf59

Observation aaf8c749-40c7-4757-a814-50bc26cf75e1 · outbound

This paper cites RedAgent: Red Teaming Large Language Models with Context-aware Autonomous Language Agent.

Exploring the Vulnerability of the Content Moderation Guardrail in Large Language Models via Intent Manipulation RedAgent: Red Teaming Large Language Models with Context-aware Autonomous Language Agent

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T14:32:30.773525Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:32:30.773525Z digest=sha256:974d164f848e848cc233f4fea4239ac052e84d30cb48968708fb68abbb2a77aa

Observation 739e6a25-7f9e-4c5f-9925-4b25d9d16e52 · outbound

This paper cites an unresolved cited work.

Exploring the Vulnerability of the Content Moderation Guardrail in Large Language Models via Intent Manipulation Unresolved cited work

Reference 38

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:32:33.698410Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T14:32:30.817410Z digest=sha256:5c2c9a65e723f64f08776e131160a3da3f51ad523125514e65a9875d11829cda

Observation 213c098c-63ae-4d11-aee0-962ef541656e · outbound

This paper cites an unresolved cited work.

Exploring the Vulnerability of the Content Moderation Guardrail in Large Language Models via Intent Manipulation Unresolved cited work

Reference 39

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:32:33.523536Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T14:32:30.859591Z digest=sha256:2e5b3a43679054d551edbb8f3b33a94be7a880800bdb38cb47ae3842d421d7a4

Observation 791bdf0f-7e6b-4e1e-bcd8-d243c270fae0 · outbound

This paper cites an unresolved cited work.

Exploring the Vulnerability of the Content Moderation Guardrail in Large Language Models via Intent Manipulation Unresolved cited work

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T14:32:30.919666Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:32:30.919666Z digest=sha256:176b1ffef6d507a9cc408c34090bc1a4f2e74a19659ea76269cf3fb45047c798

Observation e2fcdcaf-b1b8-419a-b03b-8dc884c67fa1 · outbound

This paper cites Jailbreak Attacks and Defenses Against Large Language Models: A Survey.

Exploring the Vulnerability of the Content Moderation Guardrail in Large Language Models via Intent Manipulation Jailbreak Attacks and Defenses Against Large Language Models: A Survey

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T14:32:31.034762Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:32:31.034762Z digest=sha256:bea46196cebdd9ba607d4f5afbcfa8e4e35db7f19374d3ab5ae436a46afc3c78

Observation 041362cf-344d-431d-9f65-b2f644b51495 · outbound

This paper cites an unresolved cited work.

Exploring the Vulnerability of the Content Moderation Guardrail in Large Language Models via Intent Manipulation Unresolved cited work

Reference 42

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:32:33.326229Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T14:32:31.088188Z digest=sha256:fc2e7c59e2478915f8fc80591c2ed2058da67c3bc0683dc4023e7acb4c07855d

Observation 4cf9ed77-4cdd-4c78-ba1c-76616ac9bf66 · outbound

This paper cites GPTFUZZER: Red Teaming Large Language Models with Auto-Generated Jailbreak Prompts.

Exploring the Vulnerability of the Content Moderation Guardrail in Large Language Models via Intent Manipulation GPTFUZZER: Red Teaming Large Language Models with Auto-Generated Jailbreak Prompts

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T14:32:31.124752Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:32:31.124752Z digest=sha256:ca3da5b9662924e0fbca62dd4a93915fc383e66ae03f1b6d0e4876e5f33b4a81

Observation deb7a500-333c-4e9e-af3a-52a49d38e0eb · outbound

This paper cites an unresolved cited work.

Exploring the Vulnerability of the Content Moderation Guardrail in Large Language Models via Intent Manipulation Unresolved cited work

Reference 44

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:32:33.130348Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T14:32:31.185153Z digest=sha256:84119ccade2cd7354a7f41dc1c7af45afc47d673ff123992bbd808240f46c238

Observation 7549e762-cf0c-45fa-9cd5-6c51e9a911ec · outbound

This paper cites an unresolved cited work.

Exploring the Vulnerability of the Content Moderation Guardrail in Large Language Models via Intent Manipulation Unresolved cited work

Reference 45

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:32:33.012359Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T14:32:31.272832Z digest=sha256:11a00e66518be52d49936b48bcdd23431424b46061f2edb99ac1853ae90cb9c2

Observation cafffd9b-ee29-49b4-81c8-a09b699277a3 · outbound

This paper cites an unresolved cited work.

Exploring the Vulnerability of the Content Moderation Guardrail in Large Language Models via Intent Manipulation Unresolved cited work

Reference 46

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:32:32.862069Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T14:32:31.346209Z digest=sha256:f18f7a7b083d5e8e8659eef4b6d859eaac79adc4cb9100b219c9fb14886362fe

Observation 306dfaa0-926f-4e67-9edb-c8e95cb94cd8 · outbound

This paper cites WordGame: Efficient & Effective LLM Jailbreak via Simultaneous Obfuscation in Query and Response.

Exploring the Vulnerability of the Content Moderation Guardrail in Large Language Models via Intent Manipulation WordGame: Efficient & Effective LLM Jailbreak via Simultaneous Obfuscation in Query and Response

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T14:32:31.393629Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:32:31.393629Z digest=sha256:fc2e2849b7c96dae9f7ece0cacc8c94bd9571f4cc6472a3249050fbf7ae61cc8

Observation 363920a5-7472-43d6-bfb9-c5b14637ff60 · outbound

This paper cites an unresolved cited work.

Exploring the Vulnerability of the Content Moderation Guardrail in Large Language Models via Intent Manipulation Unresolved cited work

Reference 48

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:32:32.706784Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T14:32:31.465433Z digest=sha256:9bce4974d4e7b884a25a47a211037649c2baa33a4a487194ecbca5d423ea455a

Observation 8c3068a1-aab5-4bc1-8738-4a37c628852f · outbound

This paper cites Autodan: Interpretable gradient-based adversarial attacks on large language models.

Exploring the Vulnerability of the Content Moderation Guardrail in Large Language Models via Intent Manipulation Autodan: Interpretable gradient-based adversarial attacks on large language models

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:32:32.486577Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T14:32:31.558726Z digest=sha256:8bfec9fd293fa661ff63e5c4c76223f5e27af4db882b68e206d9dd6e397c9a18

Observation 06bbe084-a062-490a-81fd-b9a1ac80b7ae · outbound

This paper cites Universal and Transferable Adversarial Attacks on Aligned Language Models.

Exploring the Vulnerability of the Content Moderation Guardrail in Large Language Models via Intent Manipulation Universal and Transferable Adversarial Attacks on Aligned Language Models

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T14:32:31.644019Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:32:31.644019Z digest=sha256:c415b55381c76e8990c9fcd04a251068aeba5d8fdeb65fd815730b71c9d165da

Pith citing papers

Observation 4a48bd81-4ad5-4d9e-b690-0bbcb0aab7c9 · inbound

Learning to Conceal Risk: Controllable Multi-turn Red Teaming for LLMs in the Financial Domain cites this paper.

Learning to Conceal Risk: Controllable Multi-turn Red Teaming for LLMs in the Financial Domain Exploring the Vulnerability of the Content Moderation Guardrail in Large Language Models via Intent Manipulation

Reference 43

Resolution
verified exact
arxiv_id, observed 2026-05-18T17:51:41.887953Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-18T17:49:42.112564Z digest=sha256:70563c38c6192db182f3103c1688bfaebd519a24bedb74a88607728b9830a6a5

Observation 49805767-a29c-4379-acc2-cebec45087db · inbound

Break Me If You Can: Self-Jailbreaking of Aligned LLMs via Lexical Insertion Prompting cites this paper.

Break Me If You Can: Self-Jailbreaking of Aligned LLMs via Lexical Insertion Prompting Exploring the Vulnerability of the Content Moderation Guardrail in Large Language Models via Intent Manipulation

Reference 10

Resolution
malformed identifier
arxiv_id, observed 2026-05-16T18:08:13.066346Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-16T18:04:44.543311Z digest=sha256:b7f4fc0c7c13866a2b84bfd5325c368f1ed46886923627a7cad5d794440865b1

Observation d40ec3d7-7faa-417d-b17c-e93b52d8736f · inbound

Incomplete Prompt Jailbreaks in Large Language Models cites this paper.

Incomplete Prompt Jailbreaks in Large Language Models Exploring the Vulnerability of the Content Moderation Guardrail in Large Language Models via Intent Manipulation

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-02T13:18:40.623601Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T13:18:40.623601Z digest=sha256:87fb61eadb0853ded033b3eacdcbb652770e15cb0a88ba41b645edf7f81db494