Pith. sign in

Paper Citation Record · LEDGER

Learning diverse attacks on large language models for robust red-teaming and safety tuning

As of 20 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 10 inbound Pith citation observations for arXiv:2405.18540.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2405.18540 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 10 of 10 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 10 of 10 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-10T18:36:49.713119Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-11T15:41:17.036171Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation f801d236-71e2-4429-8ac1-9d7c6e8556b0 · inbound

Can Safety Fine-Tuning Be More Principled? Lessons Learned from Cybersecurity cites this paper.

Can Safety Fine-Tuning Be More Principled? Lessons Learned from Cybersecurity Learning diverse attacks on large language models for robust red-teaming and safety tuning

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-10T18:36:49.713119Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:36:49.713119Z digest=sha256:0ff867f6ca50abe6fc03569b95ceb2a3aa528867401570e07c69e18d74a49a8a

Observation f015b579-a826-4ced-b02f-389e2d8872fc · inbound

Neural Genetic Search in Discrete Spaces cites this paper.

Neural Genetic Search in Discrete Spaces Learning diverse attacks on large language models for robust red-teaming and safety tuning

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-08T18:14:25.844619Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T18:14:25.844619Z digest=sha256:fa927a34ac9d2504d6b86206dfd0edf018f05b9bd86c30942ae4b5507c184776

Observation 9691966f-846a-41ba-8fad-cb9078025d7d · inbound

Adversarial Preference Learning for Robust LLM Alignment cites this paper.

Adversarial Preference Learning for Robust LLM Alignment Learning diverse attacks on large language models for robust red-teaming and safety tuning

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T12:35:21.669202Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:35:21.669202Z digest=sha256:8466343dd9b4788704b11ef7bbea5ce72371a4720ecbdc100e283b3b4d735509

Observation 8b7839bf-1144-4e59-9590-e49c9f2bbb38 · inbound

Jailbreak-R1: Exploring the Jailbreak Capabilities of LLMs via Reinforcement Learning cites this paper.

Jailbreak-R1: Exploring the Jailbreak Capabilities of LLMs via Reinforcement Learning Learning diverse attacks on large language models for robust red-teaming and safety tuning

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T12:01:08.855205Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:01:08.855205Z digest=sha256:14763627f07c92be97ea300ce4c105343c0ca909279c0ec2cb784946eb2e016d

Observation 26fcdc41-bd95-4bb1-b0d9-aef4fda1bea2 · inbound

Watch, Listen, Understand, Mislead: Tri-modal Adversarial Attacks on Short Videos for Content Appropriateness Evaluation cites this paper.

Watch, Listen, Understand, Mislead: Tri-modal Adversarial Attacks on Short Videos for Content Appropriateness Evaluation Learning diverse attacks on large language models for robust red-teaming and safety tuning

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T17:09:22.585003Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:09:22.585003Z digest=sha256:684bce80322fdef1b8ae26712625a1f8e61d9c94b7c5441df8b3bf6a5b3e8863

Observation a1881534-8996-4c2b-9edf-fed53cb73eb8 · inbound

From Seed to Harvest: Augmenting Human Creativity with AI for Red-teaming Text-to-Image Models cites this paper.

From Seed to Harvest: Augmenting Human Creativity with AI for Red-teaming Text-to-Image Models Learning diverse attacks on large language models for robust red-teaming and safety tuning

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T14:45:20.466196Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:45:20.466196Z digest=sha256:b5c43120342510941fb41890c7bd2d5eed4349f688605a00f7b710c0bfe53007

Observation e3620943-fba8-4cdf-8c73-9b0b3674bd14 · inbound

Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM cites this paper.

Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM Learning diverse attacks on large language models for robust red-teaming and safety tuning

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-05T23:13:03.820155Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:13:03.820155Z digest=sha256:40e57dfa7f18b4dec14b309a8a0d42a29678d7b3508c2386a95c8d9639e6dc54

Observation eec2fb5a-a4f3-44fc-835b-60ebb1957c71 · inbound

Layer-Wise Perturbations via Sparse Autoencoders for Adversarial Text Generation cites this paper.

Layer-Wise Perturbations via Sparse Autoencoders for Adversarial Text Generation Learning diverse attacks on large language models for robust red-teaming and safety tuning

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-05T20:31:36.553261Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:31:36.553261Z digest=sha256:9d24da0cda96a339343caf2178b7b916b066c5733ce1ea586a8a1fd4f36f75c6

Observation 7c358abf-7bf5-4559-8649-380cf8d6e359 · inbound

LLM in the Middle: A Systematic Review of Threats and Mitigations to Real-World LLM-based Systems cites this paper.

LLM in the Middle: A Systematic Review of Threats and Mitigations to Real-World LLM-based Systems Learning diverse attacks on large language models for robust red-teaming and safety tuning

Reference 169

Resolution
unresolved
no resolver link, observed 2026-08-04T17:46:19.425971Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T17:46:19.425971Z digest=sha256:ca6a2442d3ee7d6fdd24cb5e3859d5b64acb312de7cb83c8b1c3e3ebf46a31ba

Observation 305359b0-5c56-4033-a8ad-0c73cf638e2f · inbound

Stable-GFlowNet: Toward Diverse and Robust LLM Red-Teaming via Contrastive Trajectory Balance cites this paper.

Stable-GFlowNet: Toward Diverse and Robust LLM Red-Teaming via Contrastive Trajectory Balance Learning diverse attacks on large language models for robust red-teaming and safety tuning

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-11T15:41:17.115469Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-05-09T19:29:36.348680Z digest=sha256:da9e44c9648f6819a3687fc2e83a1c2282030cbc0018512f17d6bd607bb5016f