Pith. sign in

Paper Citation Record · LEDGER

Explore, Establish, Exploit: Red Teaming Language Models from Scratch

As of 7 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 29 inbound Pith citation observations for arXiv:2306.09442.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2306.09442 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 29 of 29 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 29 of 29 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T12:35:20.649185Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

7
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation e6f0ac5f-7df8-4bf4-a62d-954b350d54a1 · inbound

A Comprehensive Overview of Large Language Models cites this paper.

A Comprehensive Overview of Large Language Models Explore, Establish, Exploit: Red Teaming Language Models from Scratch

Reference 179

Resolution
verified exact
arxiv_id, observed 2026-05-19T20:28:39.570784Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-19T20:28:38.900026Z digest=sha256:a9c73c40eb1f41ed9e6c3da4f5e0edf436422dd310a792a7c40062822acdc4ee

Observation 21b70197-ee2c-45d3-97f5-10495b97384d · inbound

Baseline Defenses for Adversarial Attacks Against Aligned Language Models cites this paper.

Baseline Defenses for Adversarial Attacks Against Aligned Language Models Explore, Establish, Exploit: Red Teaming Language Models from Scratch

Reference 8

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T23:24:40.070684Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-13T23:24:39.835347Z digest=sha256:8708240e862a85798a3d7920dfbfafce80ede6caeae683bc3da9734dc2238e54

Observation 7f4390b7-62ee-445a-8216-340d244f8201 · inbound

Catastrophic Jailbreak of Open-source LLMs via Exploiting Generation cites this paper.

Catastrophic Jailbreak of Open-source LLMs via Exploiting Generation Explore, Establish, Exploit: Red Teaming Language Models from Scratch

Reference 5

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T22:00:51.519583Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-16T22:00:51.487120Z digest=sha256:0ece5e27862d8e792f7cd5da227b1899e963e89e66d33263ae58000fee36311a

Observation d6ab93bd-c519-45f4-a798-b834acb4f06f · inbound

Agent AI: Surveying the Horizons of Multimodal Interaction cites this paper.

Agent AI: Surveying the Horizons of Multimodal Interaction Explore, Establish, Exploit: Red Teaming Language Models from Scratch

Reference 31

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T14:25:59.177956Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-18T14:25:58.876978Z digest=sha256:442eb9a5e36e2d4340d007b22d77672f312c9645e742cddbab752d783d922be2

Observation febf8415-c6bf-42ff-b85c-77200f92c9da · inbound

TrustLLM: Trustworthiness in Large Language Models cites this paper.

TrustLLM: Trustworthiness in Large Language Models Explore, Establish, Exploit: Red Teaming Language Models from Scratch

Reference 231

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T11:17:08.665940Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-18T11:17:08.108565Z digest=sha256:92784e2444b764c9851c10dec443833e2697feec3fbf51634a541c8d5626e6c0

Observation 3eb0970a-203c-4127-8961-9e16f4df2c85 · inbound

Jailbreak Attacks and Defenses Against Large Language Models: A Survey cites this paper.

Jailbreak Attacks and Defenses Against Large Language Models: A Survey Explore, Establish, Exploit: Red Teaming Language Models from Scratch

Reference 12

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T02:20:44.567150Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T02:20:44.368219Z digest=sha256:ccc644f94296e7ae55b3669b31aa29d00d8c931c6c0d2f29933af750fab13779

Observation 9b32b33d-7756-4ec4-913e-7c6d1be463bd · inbound

Adversarial Preference Learning for Robust LLM Alignment cites this paper.

Adversarial Preference Learning for Robust LLM Alignment Explore, Establish, Exploit: Red Teaming Language Models from Scratch

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T12:35:20.649185Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:35:20.649185Z digest=sha256:be76790448bd15d1d057d4b54228642c97ca9cedf589336092d7b636297fe455

Observation ae34b8d3-fbac-43ec-a207-7dba2ef5b36d · inbound

RedRFT: A Light-Weight Benchmark for Reinforcement Fine-Tuning-Based Red Teaming cites this paper.

RedRFT: A Light-Weight Benchmark for Reinforcement Fine-Tuning-Based Red Teaming Explore, Establish, Exploit: Red Teaming Language Models from Scratch

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T10:55:20.057855Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:55:20.057855Z digest=sha256:b0effb200c0707563e25d48cdfe2624c3a71b60b5b9390b39c5ca87410b669b7

Observation cfba1cf5-35fd-49e4-81d5-182a7d296090 · inbound

FORTRESS: Frontier Risk Evaluation for National Security and Public Safety cites this paper.

FORTRESS: Frontier Risk Evaluation for National Security and Public Safety Explore, Establish, Exploit: Red Teaming Language Models from Scratch

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T00:14:59.966022Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:14:59.966022Z digest=sha256:edb6dea28b98255a67f007ecd033a897de0b12711b3dedf52acc26dd8f773dbe

Observation d1c03a4e-92e4-4a51-9c7c-61118fc31784 · inbound

Kaleidoscopic Teaming in Multi Agent Simulations cites this paper.

Kaleidoscopic Teaming in Multi Agent Simulations Explore, Establish, Exploit: Red Teaming Language Models from Scratch

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T23:35:47.454816Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:35:47.454816Z digest=sha256:be17a55dd407740187582505b6ea4c5b78fd275dbafbe6b4ae809e6a44b7e6a9

Observation 54a66195-b915-43e4-9959-84fd95786d5c · inbound

SoK: A Comprehensive Security Analysis of Jailbreak Resilience in GPT and DeepSeek Models cites this paper.

SoK: A Comprehensive Security Analysis of Jailbreak Resilience in GPT and DeepSeek Models Explore, Establish, Exploit: Red Teaming Language Models from Scratch

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T23:20:55.144189Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:20:55.144189Z digest=sha256:9512aaa4ce0ada8a577d269ffcd3cc4388a1ac8977ccccc3ab6e1ad288b2139a

Observation ee47c9ff-dcb0-43a9-a097-9fad4e0e75d8 · inbound

PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training cites this paper.

PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training Explore, Establish, Exploit: Red Teaming Language Models from Scratch

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T17:35:48.010807Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:35:48.010807Z digest=sha256:97a5528807bbcc56ef3f0a3209ece747598b539efca2951460f2edf356e78dd7

Observation 8fa9f4d3-d632-40d9-a50d-d4ed6826d947 · inbound

Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM cites this paper.

Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM Explore, Establish, Exploit: Red Teaming Language Models from Scratch

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-05T23:13:03.550174Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:13:03.550174Z digest=sha256:90e9a298271e95c1ee43c1208e7733bc18b8208f6fb69ddb5c30776a7ea28abf

Observation 532f346b-bf62-476d-9d95-24477cb4d118 · inbound

Quantized but Deceptive? A Multi-Dimensional Truthfulness Evaluation of Quantized LLMs cites this paper.

Quantized but Deceptive? A Multi-Dimensional Truthfulness Evaluation of Quantized LLMs Explore, Establish, Exploit: Red Teaming Language Models from Scratch

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-05T15:52:44.872573Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T15:52:44.872573Z digest=sha256:6be7b3b2232cf22f10fa86fe68a0a8fd249669ea2fb5d987756ae1660060cdc6

Observation eb7228e8-a91f-4b61-ae2a-d1c55a42c3b9 · inbound

Anchoring Refusal Direction: Mitigating Safety Risks in Tuning via Projection Constraint cites this paper.

Anchoring Refusal Direction: Mitigating Safety Risks in Tuning via Projection Constraint Explore, Establish, Exploit: Red Teaming Language Models from Scratch

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-04T23:10:21.273056Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:10:21.273056Z digest=sha256:445a4e4364a2b2c004549782b3f9c52e0265150ea6bc56d2b065934ebb74a260

Observation 5f56caf1-de30-425d-a0ec-e7d7fabde360 · inbound

MoGU V2: Toward a Higher Pareto Frontier Between Model Usability and Security cites this paper.

MoGU V2: Toward a Higher Pareto Frontier Between Model Usability and Security Explore, Establish, Exploit: Red Teaming Language Models from Scratch

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-04T23:09:41.508015Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:09:41.508015Z digest=sha256:a955ec0c83d76266ef6a2dc8f13c0ce94b9341a0ff8c89dc49033fe310437a61

Observation 0cf155fc-6461-4c27-a3bc-1c990bfdb00c · inbound

Tailored untruths: How personalisation challenges LLM safeguards cites this paper.

Tailored untruths: How personalisation challenges LLM safeguards Explore, Establish, Exploit: Red Teaming Language Models from Scratch

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-04T09:55:53.355255Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T09:55:53.355255Z digest=sha256:b56e7ff982c3692d3cc2c60cbff19ebc31fa3ced23da22247908f7411af3fbb7

Observation ab94034e-cb58-4d6f-acc7-7f9ffdad0f68 · inbound

Learning Uncertainty from Sequential Internal Dispersion in Large Language Models cites this paper.

Learning Uncertainty from Sequential Internal Dispersion in Large Language Models Explore, Establish, Exploit: Red Teaming Language Models from Scratch

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-10T08:48:02.413805Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-10T08:36:39.242766Z digest=sha256:599650877828a381228cfc7406199b3eb5e37adccf75ed61e45c081ec3b819e0

Observation d8955db2-bcdd-499c-b8ec-3b9aad7f8b3e · inbound

Reverse Constitutional AI: A Framework for Controllable Toxic Data Generation via Probability-Clamped RLAIF cites this paper.

Reverse Constitutional AI: A Framework for Controllable Toxic Data Generation via Probability-Clamped RLAIF Explore, Establish, Exploit: Red Teaming Language Models from Scratch

Reference 45

Resolution
verified exact
arxiv_id, observed 2026-05-10T12:15:22.405159Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-10T04:35:51.223025Z digest=sha256:d250ac88348d788b554a0127f3a4f1be7b4f0e404110b66e8d35a5b7269696d4

Observation 27e374e1-026f-4884-9734-44dca0ac509e · inbound

A Systematic Investigation of RL-Jailbreaking in LLMs cites this paper.

A Systematic Investigation of RL-Jailbreaking in LLMs Explore, Establish, Exploit: Red Teaming Language Models from Scratch

Reference 1

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T01:40:52.596590Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-11T01:32:42.151644Z digest=sha256:789cb03a34de8d0bf8b44c55fc5d48463307a4c7d5d07b05d2c1dc8af6fecdd8

Observation 1c8e15fe-8f07-42ae-8d0c-98f83597482b · inbound

A Systematic Investigation of RL-Jailbreaking in LLMs cites this paper.

A Systematic Investigation of RL-Jailbreaking in LLMs Explore, Establish, Exploit: Red Teaming Language Models from Scratch

Reference 1

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T13:35:46.533799Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-30T22:59:07.861941Z digest=sha256:01d240a2dceb78fb8f2f5bc8f5f21657d626271065e3db23e37e2e523acfab53

Observation 013aef7e-a486-4a92-b0c6-1e03faf6af9e · inbound

A Systematic Investigation of RL-Jailbreaking in LLMs cites this paper.

A Systematic Investigation of RL-Jailbreaking in LLMs Explore, Establish, Exploit: Red Teaming Language Models from Scratch

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-02T14:41:05.649595Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T14:41:05.649595Z digest=sha256:0eb778593e93d98597d78e7da5b69596f574f289f5a860ee0b68d25ffee9287c

Observation b326c0f5-09b8-470d-bb57-d1c0b650a771 · inbound

Correcting Influence: Unboxing LLM Outputs with Orthogonal Latent Spaces cites this paper.

Correcting Influence: Unboxing LLM Outputs with Orthogonal Latent Spaces Explore, Establish, Exploit: Red Teaming Language Models from Scratch

Reference 207

Resolution
metadata mismatch
arxiv_id, observed 2026-05-14T20:17:54.348080Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-14T20:17:01.224864Z digest=sha256:59f48b8b73cd854ac106c7e6d9c45dbeaf0c6009c7e576fd9aba2cba3db5160e

Observation d7bd52b0-f8a0-451f-9ea4-8ab3a91f5838 · inbound

Formal Methods Meet LLMs: Auditing, Monitoring, and Intervention for Compliance of Advanced AI Systems cites this paper.

Formal Methods Meet LLMs: Auditing, Monitoring, and Intervention for Compliance of Advanced AI Systems Explore, Establish, Exploit: Red Teaming Language Models from Scratch

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-05-20T17:28:48.020543Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-20T17:26:02.346222Z digest=sha256:ffac0d59602ecb3824eff88e6ca34f417c8e1f1aa7da43e11534f44ae6aef895

Observation 622045b0-0c14-4d76-a948-691036c06a9d · inbound

Boosting Self-Consistency with Ranking cites this paper.

Boosting Self-Consistency with Ranking Explore, Establish, Exploit: Red Teaming Language Models from Scratch

Reference 136

Resolution
metadata mismatch
arxiv_id, observed 2026-06-28T06:51:44.267655Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-28T06:49:58.051659Z digest=sha256:7d0a6d2b6c7e1a474cf751db7fe49dea568cd518f53fb3106e4b6af2fc39e39e

Observation c911afea-adfd-4ed9-aa69-f2d0abe6abf6 · inbound

Data Selection Through Iterative Self-Filtering for Vision-Language Settings cites this paper.

Data Selection Through Iterative Self-Filtering for Vision-Language Settings Explore, Establish, Exploit: Red Teaming Language Models from Scratch

Reference 43

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T09:49:44.864881Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-26T09:22:47.537137Z digest=sha256:2e06b74c64d1ec5fd64fccf98fbf1c2cf6579dfe0f5a15881c159c5de818a115

Observation 1c5b64cb-10fe-4708-8cb7-f0e89a5b8e2a · inbound

On-Policy Self-Distillation with Sampled Demonstrations Reduces Output Diversity cites this paper.

On-Policy Self-Distillation with Sampled Demonstrations Reduces Output Diversity Explore, Establish, Exploit: Red Teaming Language Models from Scratch

Reference 29

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T21:00:08.519430Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-25T19:23:56.452083Z digest=sha256:d9282b10a0385181e2c4c2090ea604f178da50e25c6efe9a4dc14f954110f1d1

Observation c257477e-7e72-4329-a983-e8ba226c1af0 · inbound

Exposure is not manifestation: measurement target and output resolution jointly determine which behavioural-faithfulness evaluator wins cites this paper.

Exposure is not manifestation: measurement target and output resolution jointly determine which behavioural-faithfulness evaluator wins Explore, Establish, Exploit: Red Teaming Language Models from Scratch

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-02T07:44:08.962913Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T07:44:08.962913Z digest=sha256:587985a31e6adb6967b8e863f127a3a080edd536326d34165546b3e04530d7ad

Observation 40e14c18-6f9e-4d12-aa77-8bfc6d0d6d39 · inbound

Just Keep Prompting: Evaluating Repetitive Socratic Prompting in VLMs cites this paper.

Just Keep Prompting: Evaluating Repetitive Socratic Prompting in VLMs Explore, Establish, Exploit: Red Teaming Language Models from Scratch

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-02T15:07:17.112776Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T15:07:17.112776Z digest=sha256:755594fc8accf00ef3f126acc434d99253a61e3b183fdf4b6af82e8be83b8d1d