Pith. sign in

Paper Citation Record · LEDGER

Attack Prompt Generation for Red Teaming and Defending Large Language Models

As of 7 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 8 inbound Pith citation observations for arXiv:2310.12505.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2310.12505 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 8 of 8 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 8 of 8 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T12:35:20.826210Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-18T05:25:54.592662Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 15d46676-6f82-46d9-8fe3-7b01c8c96701 · inbound

Adversarial Preference Learning for Robust LLM Alignment cites this paper.

Adversarial Preference Learning for Robust LLM Alignment Attack Prompt Generation for Red Teaming and Defending Large Language Models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T12:35:20.826210Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:35:20.826210Z digest=sha256:bb37b87a17e6a4e118493c89a829163e7a9136d7ddc54b1cae3c7d172132223d

Observation b06c2a86-7fd2-4537-92e4-511cf268330d · inbound

RedRFT: A Light-Weight Benchmark for Reinforcement Fine-Tuning-Based Red Teaming cites this paper.

RedRFT: A Light-Weight Benchmark for Reinforcement Fine-Tuning-Based Red Teaming Attack Prompt Generation for Red Teaming and Defending Large Language Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T10:55:20.050296Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:55:20.050296Z digest=sha256:82e62cac9e2a0b80c27547a8a0693984d7e2b885c4cfe8cf2f9ff7dbd85f66e7

Observation 2e56ca2d-ef73-4583-ba60-800825df4643 · inbound

GuardVal: Dynamic Large Language Model Jailbreak Evaluation for Comprehensive Safety Testing cites this paper.

GuardVal: Dynamic Large Language Model Jailbreak Evaluation for Comprehensive Safety Testing Attack Prompt Generation for Red Teaming and Defending Large Language Models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T18:40:52.210500Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:40:52.210500Z digest=sha256:66c990337adb75cc82f6cf7fd005fac39be1fa48a6105bd564969e60b9ba3cd4

Observation cbef038d-dd7b-4ab2-abae-e276cb072e90 · inbound

Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM cites this paper.

Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM Attack Prompt Generation for Red Teaming and Defending Large Language Models

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-05T23:13:02.478257Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:13:02.478257Z digest=sha256:3e67c8d1f826f369bf8515ca0dc5ee694d58fdf486ac28580fd7924d46a3f618

Observation 3409c0bc-093a-48b3-8b4c-02487b8338b1 · inbound

Anchoring Refusal Direction: Mitigating Safety Risks in Tuning via Projection Constraint cites this paper.

Anchoring Refusal Direction: Mitigating Safety Risks in Tuning via Projection Constraint Attack Prompt Generation for Red Teaming and Defending Large Language Models

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-04T23:10:21.359755Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:10:21.359755Z digest=sha256:915ec4b6f1b583196acdd8ecd1b10a0f371b67478c13c5c6cac1004315000efa

Observation bdcdeac0-827d-4bbc-9058-f8fe59b29b4a · inbound

SAID: Safety-Aware Intent Defense via Prefix Probing for Large Language Models cites this paper.

SAID: Safety-Aware Intent Defense via Prefix Probing for Large Language Models Attack Prompt Generation for Red Teaming and Defending Large Language Models

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-18T05:25:54.594563Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-18T05:22:30.509147Z digest=sha256:ac0e3ffe69cbdfc12ad1b8c5c8b7465931d7ec40e1f7d83caa6d4596882ec5cb

Observation cf3b0c90-f06b-44cb-ba23-b8b51fe25423 · inbound

Beyond Context: Large Language Models' Failure to Grasp Users' Intent cites this paper.

Beyond Context: Large Language Models' Failure to Grasp Users' Intent Attack Prompt Generation for Red Teaming and Defending Large Language Models

Reference 54

Resolution
verified exact
arxiv_id, observed 2026-05-16T20:11:13.577168Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-16T20:09:25.827452Z digest=sha256:9d941dcb04dddf057edfe4b6d083b19eaa523277c7ed1b1b6dff52e0308d3d9d

Observation 72a24dfa-255a-4af3-803c-8c8f48093676 · inbound

Adaptive Instruction Composition for Automated LLM Red-Teaming cites this paper.

Adaptive Instruction Composition for Automated LLM Red-Teaming Attack Prompt Generation for Red Teaming and Defending Large Language Models

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-11T14:06:03.793513Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-09T23:34:36.280092Z digest=sha256:b1bb4e9ce7bc4f115ae757ebfd976ba185026092c6647842dd4e530b48bd17e7