Pith. sign in

Paper Citation Record · LEDGER

Do-Not-Answer: A Dataset for Evaluating Safeguards in LLMs

As of 10 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 39 inbound Pith citation observations for arXiv:2308.13387.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2308.13387 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 39 of 39 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 39 of 39 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-09T13:14:34.116590Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

4
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 49da06db-1474-4a09-96ad-1dec12bda0b5 · inbound

Low-Resource Languages Jailbreak GPT-4 cites this paper.

Low-Resource Languages Jailbreak GPT-4 Do-Not-Answer: A Dataset for Evaluating Safeguards in LLMs

Reference 51

Resolution
metadata mismatch
arxiv_id, observed 2026-05-17T09:24:14.124720Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-17T09:24:13.911401Z digest=sha256:d2519b09e2080486c600d3825daa9b8263e0d89aca1ec48cea3fa1d3756139f5

Observation d214990f-e87f-4a7c-bd7c-551b0ac587d2 · inbound

DeepSeek LLM: Scaling Open-Source Language Models with Longtermism cites this paper.

DeepSeek LLM: Scaling Open-Source Language Models with Longtermism Do-Not-Answer: A Dataset for Evaluating Safeguards in LLMs

Reference 151

Resolution
verified exact
arxiv_id, observed 2026-05-11T06:08:05.807659Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-11T06:08:05.550346Z digest=sha256:ddbb397856ec4c9e70a77061e6dbde06e969a068a80c2a138529c216f77c2273

Observation f80178cc-aec5-402b-b6c5-a92bbc41ba05 · inbound

TrustLLM: Trustworthiness in Large Language Models cites this paper.

TrustLLM: Trustworthiness in Large Language Models Do-Not-Answer: A Dataset for Evaluating Safeguards in LLMs

Reference 73

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T11:17:08.446480Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-18T11:17:08.108565Z digest=sha256:c1c1a1cd91caed8949623cbbaa8a072bb3e4df26b15526bc2f8a0c3da65cedff

Observation dbae861b-3cd9-4b8d-ad84-67e64f05caba · inbound

DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model cites this paper.

DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model Do-Not-Answer: A Dataset for Evaluating Safeguards in LLMs

Reference 149

Resolution
verified exact
arxiv_id, observed 2026-05-11T05:36:26.527158Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-11T05:36:26.207359Z digest=sha256:322974afcf972a6a0002f6a0baf17096a8438ccec797b6c0a92bc7233f04f7fe

Observation 4b290f1a-438c-44dc-a7ec-95464f90b851 · inbound

Jailbreak Attacks and Defenses Against Large Language Models: A Survey cites this paper.

Jailbreak Attacks and Defenses Against Large Language Models: A Survey Do-Not-Answer: A Dataset for Evaluating Safeguards in LLMs

Reference 96

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T02:20:44.861198Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-15T02:20:44.368219Z digest=sha256:06a86486a913ef79fbdd7f9dc7b9995f0ef1675453838f9b5b450180976cdc55

Observation aebf8f1a-3c1c-4fe8-a5c7-c90c725ec6d2 · inbound

Vulnerability Mitigation for Safety-Aligned Language Models via Debiasing cites this paper.

Vulnerability Mitigation for Safety-Aligned Language Models via Debiasing Do-Not-Answer: A Dataset for Evaluating Safeguards in LLMs

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-09T13:14:34.116590Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T13:14:34.116590Z digest=sha256:ab691083d954c5d11f8135e153bbc79580a39f7a1bdc091505f1c7d3d678d740

Observation 025fa25f-3b8a-4744-8796-914ab383d9b6 · inbound

Safety Reasoning with Guidelines cites this paper.

Safety Reasoning with Guidelines Do-Not-Answer: A Dataset for Evaluating Safeguards in LLMs

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-08T23:50:36.181837Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T23:50:36.181837Z digest=sha256:7ac8745f167d28488d2faee8b8f801a63beb8152640d1171e673ad28e0733649

Observation 8dea832e-a657-4a75-9287-1323bcb16dd5 · inbound

SweEval: Do LLMs Really Swear? A Safety Benchmark for Testing Limits for Enterprise Use cites this paper.

SweEval: Do LLMs Really Swear? A Safety Benchmark for Testing Limits for Enterprise Use Do-Not-Answer: A Dataset for Evaluating Safeguards in LLMs

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T14:52:30.727155Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:52:30.727155Z digest=sha256:a1d92261f55b3e7d424e7cc37241efb4e69a64a2916172b361bdc2869e2b505e

Observation b4e6b0ea-e62c-4c59-a12d-910c41a6b0cb · inbound

Xinyu AI Search: Enhanced Relevance and Comprehensive Results with Rich Answer Presentations cites this paper.

Xinyu AI Search: Enhanced Relevance and Comprehensive Results with Rich Answer Presentations Do-Not-Answer: A Dataset for Evaluating Safeguards in LLMs

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-07T13:26:10.644766Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:26:10.644766Z digest=sha256:a1751657865cd4f2cdb1b822d8e2155fdbf760934d8abb37f2dcf196343a1ded

Observation 2d3a899b-48da-4516-b36b-31962dfbab41 · inbound

Beyond the Surface: Measuring Self-Preference in LLM Judgments cites this paper.

Beyond the Surface: Measuring Self-Preference in LLM Judgments Do-Not-Answer: A Dataset for Evaluating Safeguards in LLMs

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T11:26:06.467953Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:26:06.467953Z digest=sha256:58679436cdef6e143e240cc40ea29e1af232e6fd203a8da2739c422cd12a21b1

Observation a77cdee8-0522-41fa-84d8-875e81486a24 · inbound

Does Math Reasoning Improve General LLM Capabilities? Understanding Transferability of LLM Reasoning cites this paper.

Does Math Reasoning Improve General LLM Capabilities? Understanding Transferability of LLM Reasoning Do-Not-Answer: A Dataset for Evaluating Safeguards in LLMs

Reference 234

Resolution
verified exact
arxiv_id, observed 2026-05-19T01:01:09.979243Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-19T01:01:09.840919Z digest=sha256:9736ce942b8d67e505ce18c6b6f57875669707f43ef9ac7900abce942a16bae7

Observation 0c3a4973-30b9-47b5-9a48-97af1325e9d3 · inbound

SEALGuard: Safeguarding the Multilingual Conversations in Southeast Asian Languages for LLM Software Systems cites this paper.

SEALGuard: Safeguarding the Multilingual Conversations in Southeast Asian Languages for LLM Software Systems Do-Not-Answer: A Dataset for Evaluating Safeguards in LLMs

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-06T18:28:45.637554Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:28:45.637554Z digest=sha256:64d3144208c68fd552dba94a5883cc337daf2c44076b288b132e8987485b500c

Observation 6e3d5fd9-d925-426f-aeed-56809efd5295 · inbound

Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM cites this paper.

Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM Do-Not-Answer: A Dataset for Evaluating Safeguards in LLMs

Reference 124

Resolution
unresolved
no resolver link, observed 2026-08-05T23:13:04.544081Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:13:04.544081Z digest=sha256:0ce97458623890bd033daaef3828698927e77a00401725ac35ab953aaac5fd08

Observation dde2c32f-5be4-4115-a64b-ff810cd02d84 · inbound

Revisiting Robustness for LLM Safety Alignment via Selective Geometry Control cites this paper.

Revisiting Robustness for LLM Safety Alignment via Selective Geometry Control Do-Not-Answer: A Dataset for Evaluating Safeguards in LLMs

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-22T11:21:28.969308Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-22T11:17:03.104902Z digest=sha256:01fa65a4bb0a358fe731f44c1b084ae917bc653a8fbda54fc5e40018a3108e00

Observation aff94550-37b0-4bdf-8fcb-7c74bfd3e047 · inbound

Cooking Up Risks: Benchmarking and Reducing Food Safety Risks in Large Language Models cites this paper.

Cooking Up Risks: Benchmarking and Reducing Food Safety Risks in Large Language Models Do-Not-Answer: A Dataset for Evaluating Safeguards in LLMs

Reference 31

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T22:03:20.386027Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-13T22:02:12.555644Z digest=sha256:3a34e24d11a6aaa1c0189caeb2aab7e432375d6a086bdbdfeb825fdaad69cce0

Observation c4abdfe4-cb30-48df-a7cd-305efd06f8d7 · inbound

Guardian-as-an-Advisor: Advancing Next-Generation Guardian Models for Trustworthy LLMs cites this paper.

Guardian-as-an-Advisor: Advancing Next-Generation Guardian Models for Trustworthy LLMs Do-Not-Answer: A Dataset for Evaluating Safeguards in LLMs

Reference 80

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T06:46:36.406518Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T17:27:13.339411Z digest=sha256:a33cb8156db8b84854b0cdc0ad9021c7ce55e3e210e669a6055cca12c32a4bd8

Observation 931bb8ab-2a49-4933-bdbc-cea2a466d8e6 · inbound

The Art of (Mis)alignment: How Fine-Tuning Methods Effectively Misalign and Realign LLMs in Post-Training cites this paper.

The Art of (Mis)alignment: How Fine-Tuning Methods Effectively Misalign and Realign LLMs in Post-Training Do-Not-Answer: A Dataset for Evaluating Safeguards in LLMs

Reference 60

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T00:45:50.631891Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T18:18:56.476698Z digest=sha256:41b01508993c358a8cb46ba44bc7bb00bce50cd0184d1564ccc799827c1ce22c

Observation 329b1e73-7a7a-4401-ac4b-213d93b31f5d · inbound

Dialect vs Demographics: Quantifying LLM Bias from Implicit Linguistic Signals vs. Explicit User Profiles cites this paper.

Dialect vs Demographics: Quantifying LLM Bias from Implicit Linguistic Signals vs. Explicit User Profiles Do-Not-Answer: A Dataset for Evaluating Safeguards in LLMs

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-09T22:39:14.506405Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-09T22:35:03.952451Z digest=sha256:0235e323f5442dd6318496ac921fa810bf6197d4008187e39ccb65fb08bb4dba

Observation 67893cbf-c80c-41c7-b2d2-f3776cc73794 · inbound

MultiBreak: A Scalable and Diverse Multi-turn Jailbreak Benchmark for Evaluating LLM Safety cites this paper.

MultiBreak: A Scalable and Diverse Multi-turn Jailbreak Benchmark for Evaluating LLM Safety Do-Not-Answer: A Dataset for Evaluating Safeguards in LLMs

Reference 20

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T09:31:01.138639Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-10T16:00:32.413225Z digest=sha256:1f166539dcd0f7b6164e0022c26066973b0478db748031cdbcac6cf3d48631e0

Observation 6bacf8e7-54c3-4907-94b3-e089133574ce · inbound

A Validated Prompt Bank for Malicious Code Generation: Separating Executable Weapons from Security Knowledge in 1,554 Consensus-Labeled Prompts cites this paper.

A Validated Prompt Bank for Malicious Code Generation: Separating Executable Weapons from Security Knowledge in 1,554 Consensus-Labeled Prompts Do-Not-Answer: A Dataset for Evaluating Safeguards in LLMs

Reference 17

Resolution
metadata mismatch
arxiv_id, observed 2026-05-09T06:40:43.791322Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-08T18:11:29.066362Z digest=sha256:65f4ade05cbd61e4cf998fc925e36647555f3554184206625371172668e568a5

Observation b4835099-28c6-415c-8219-aaf1e81b83c2 · inbound

Beyond Fixed Benchmarks and Worst-Case Attacks: Dynamic Boundary Evaluation for Language Models cites this paper.

Beyond Fixed Benchmarks and Worst-Case Attacks: Dynamic Boundary Evaluation for Language Models Do-Not-Answer: A Dataset for Evaluating Safeguards in LLMs

Reference 23

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T20:06:13.694609Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-08T10:13:35.777910Z digest=sha256:391538631a9a282d4a9a1accaacaca8154a1bddeb30d9bbdf11c834c6eb277c6

Observation 8189446b-33c3-4263-aab8-515445b7585c · inbound

Navigating the Sea of LLM Evaluation: Investigating Bias in Toxicity Benchmarks cites this paper.

Navigating the Sea of LLM Evaluation: Investigating Bias in Toxicity Benchmarks Do-Not-Answer: A Dataset for Evaluating Safeguards in LLMs

Reference 30

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T05:31:23.900065Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T05:28:45.453455Z digest=sha256:587d283d954527a5c0b854d0bba66c2ed31b6caed82de8b466b7725711aa1645

Observation 1123835b-527d-428a-a568-12d987057cf5 · inbound

REFLECTOR: Internalizing Step-wise Reflection against Indirect Jailbreak cites this paper.

REFLECTOR: Internalizing Step-wise Reflection against Indirect Jailbreak Do-Not-Answer: A Dataset for Evaluating Safeguards in LLMs

Reference 19

Resolution
metadata mismatch
arxiv_id, observed 2026-05-21T06:19:42.002339Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-21T06:16:01.040236Z digest=sha256:ce940fadce2aec8ce9805186a766d59d53d66012481f19d2984b40dc590195a6

Observation 41351988-a409-43be-b92d-e6ba900d5688 · inbound

JT-SAFE-V2: Safety-by-Design Foundation Model with World-Context Data cites this paper.

JT-SAFE-V2: Safety-by-Design Foundation Model with World-Context Data Do-Not-Answer: A Dataset for Evaluating Safeguards in LLMs

Reference 16

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T13:54:44.172152Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-30T13:45:38.305767Z digest=sha256:4caef0cadb77fa008578e4208456e3fd0dbfd33bdbfe5c26ab4ef3e86ae050ad

Observation db495455-5492-4c3a-8908-647e36b023b8 · inbound

RouteScan: A Non-Intrusive Approach to Auditing MoE LLMs Safety via Expert Routing Telemetry cites this paper.

RouteScan: A Non-Intrusive Approach to Auditing MoE LLMs Safety via Expert Routing Telemetry Do-Not-Answer: A Dataset for Evaluating Safeguards in LLMs

Reference 49

Resolution
verified exact
arxiv_id, observed 2026-06-30T00:34:05.321167Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-30T00:33:35.235214Z digest=sha256:742dca2fa00fb8c29df2dc1e0e0c6dafb59d995e3285dd5a8363ad249e389385

Observation 7bafba10-78c1-43df-945c-79fc7fd30192 · inbound

IDEAFix: Evaluation Framework for Creative Defixation Prompting in LLMs cites this paper.

IDEAFix: Evaluation Framework for Creative Defixation Prompting in LLMs Do-Not-Answer: A Dataset for Evaluating Safeguards in LLMs

Reference 50

Resolution
metadata mismatch
arxiv_id, observed 2026-06-28T18:42:28.917805Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-06-28T18:42:00.845492Z digest=sha256:2edd09755300169930e7a51ccb33041855424c907c4d70c84f5127a218ef0bbf

Observation 8a5fb43b-54f1-4e5c-a1d3-f03a2f75e43e · inbound

When Behavioral Safety Evaluation Fails: A Representation-Level Perspective cites this paper.

When Behavioral Safety Evaluation Fails: A Representation-Level Perspective Do-Not-Answer: A Dataset for Evaluating Safeguards in LLMs

Reference 20

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T20:57:23.020717Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-06-27T20:04:17.744876Z digest=sha256:a3d24101c0d0adbeac822a86c9609e45d852d383a66eb8aa7fc52691f0567cfa

Observation d8086c03-e6eb-4e6f-8c34-302b925b7b0b · inbound

Unsupervised Causal Abstractions Discovery cites this paper.

Unsupervised Causal Abstractions Discovery Do-Not-Answer: A Dataset for Evaluating Safeguards in LLMs

Reference 84

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T00:59:21.203220Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-06-26T20:45:15.226889Z digest=sha256:5a8ded4da1fab557eb4db63dd342cf0635a67dc010604858fe3c28cd3fb2bc18

Observation 96db16a7-6e3b-479b-81d1-551b69abd473 · inbound

FinRED: An Expert-Guided Benchmark Generation and Evaluation Framework for Financial LLM Red-Teaming cites this paper.

FinRED: An Expert-Guided Benchmark Generation and Evaluation Framework for Financial LLM Red-Teaming Do-Not-Answer: A Dataset for Evaluating Safeguards in LLMs

Reference 26

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T04:09:34.773753Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-26T17:11:40.088809Z digest=sha256:566b0c7b577561b8d6ab2f3120740f76229379dd06798c1660df4c103a9ce4de

Observation aaa08779-a722-40ec-ae7e-ba503da2235b · inbound

Efficient Safety Benchmarking via Item Response Theory cites this paper.

Efficient Safety Benchmarking via Item Response Theory Do-Not-Answer: A Dataset for Evaluating Safeguards in LLMs

Reference 13

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T15:55:48.971641Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-07-01T15:51:00.484343Z digest=sha256:2e7c897ae0377826cb936c77b19ac3d0949845c6ce606b114b4acd63e49976d6

Observation e7a258d6-01cd-4f4c-b84e-22a31e3655f6 · inbound

Discriminatory Compliance: How LLMs Answer Queries from Protected Groups cites this paper.

Discriminatory Compliance: How LLMs Answer Queries from Protected Groups Do-Not-Answer: A Dataset for Evaluating Safeguards in LLMs

Reference 100

Resolution
verified exact
arxiv_id, observed 2026-06-26T12:59:29.165546Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-06-26T12:58:09.227648Z digest=sha256:1864d4e88912b77c90e5c81e204abbe726944a5d570f0623aa6f63c4862ae66a

Observation ea5bf22f-5b06-4ed8-9db0-30d635e15b83 · inbound

Do Encoders Suffice? A Systematic Comparison of Encoder and Decoder Safety Judges for LLM Adversarial Evaluation cites this paper.

Do Encoders Suffice? A Systematic Comparison of Encoder and Decoder Safety Judges for LLM Adversarial Evaluation Do-Not-Answer: A Dataset for Evaluating Safeguards in LLMs

Reference 22

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T20:00:07.965831Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-25T20:55:44.549142Z digest=sha256:7486bb79921bce7976552ae5ec02ce6e32f9e497f0a252044f59091565e0b8d2

Observation 6f8ddaa2-f460-4ba9-9473-dbf00df5c8e5 · inbound

Symbolic Mechanistic Data Attribution: Tracing Training Influence to Learned Behavioral Policies cites this paper.

Symbolic Mechanistic Data Attribution: Tracing Training Influence to Learned Behavioral Policies Do-Not-Answer: A Dataset for Evaluating Safeguards in LLMs

Reference 62

Resolution
verified exact
arxiv_id, observed 2026-06-30T08:04:27.971832Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-06-30T08:03:23.581020Z digest=sha256:723fc345504994f29b2e3d39a81701ff0325d4653cdf8cd0c4dadc580bd22ebe

Observation 166868af-f1aa-4a2d-bc06-e82283a56383 · inbound

BioTIER: A Refusal Benchmark for Targeted Biological Risk Mitigation cites this paper.

BioTIER: A Refusal Benchmark for Targeted Biological Risk Mitigation Do-Not-Answer: A Dataset for Evaluating Safeguards in LLMs

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-02T02:00:44.726066Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T02:00:44.726066Z digest=sha256:5cc298a489b4336554819b8e30810caa774a1ee7bdc480b71398b2ca8d03f9bc

Observation fc180841-96ec-4227-8e0e-e3a769d177a1 · inbound

Reason Before You Retrieve: Agentic Planning for Multi-modal RAG cites this paper.

Reason Before You Retrieve: Agentic Planning for Multi-modal RAG Do-Not-Answer: A Dataset for Evaluating Safeguards in LLMs

Reference 106

Resolution
unresolved
no resolver link, observed 2026-08-02T10:20:56.053655Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T10:20:56.053655Z digest=sha256:ced1e7d71527dd20456b5a4f0677d7e7391c67f4aa1b6a5815bff80e775902d8

Observation f60688c9-3ec5-4555-84be-6cf231031183 · inbound

The Mirage of LLM Guardrails: A Case Study in AI-Assisted Medical Note Manipulation cites this paper.

The Mirage of LLM Guardrails: A Case Study in AI-Assisted Medical Note Manipulation Do-Not-Answer: A Dataset for Evaluating Safeguards in LLMs

Reference 11

Resolution
unresolved
no resolver link, observed 2026-07-30T20:22:15.974831Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-30T20:22:15.974831Z digest=sha256:a83c729547baeb307e8ed9e3b04d223b9b5d8be8d22f15468f23dd65ad326b89

Observation ebea8a40-8b99-4396-8cd8-b56c0929cae9 · inbound

Forecasting Trajectory-Level Safety Risks in Black-Box Multi-Turn Interactions cites this paper.

Forecasting Trajectory-Level Safety Risks in Black-Box Multi-Turn Interactions Do-Not-Answer: A Dataset for Evaluating Safeguards in LLMs

Reference 99

Resolution
unresolved
no resolver link, observed 2026-07-30T20:11:00.919374Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-30T20:11:00.919374Z digest=sha256:bd9c2de294b55880a01b12676a57dab3711c519c65dfc0befbbbfc5131111e6f

Observation 67e98565-810a-4f74-8e17-bd34845b0412 · inbound

Chain-of-Models: Cross-Model Auditing for Bias-Robust LLM Judges cites this paper.

Chain-of-Models: Cross-Model Auditing for Bias-Robust LLM Judges Do-Not-Answer: A Dataset for Evaluating Safeguards in LLMs

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-03T00:55:21.768924Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T00:55:21.768924Z digest=sha256:7dd575db29e0ef3050ea0394f8db9a52a4dbc29d91d09dae832715e1912ab832

Observation 52acaad3-523c-437f-b54c-ec44f120bc2a · inbound

Item Response Theory for AI Safety cites this paper.

Item Response Theory for AI Safety Do-Not-Answer: A Dataset for Evaluating Safeguards in LLMs

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-06T05:39:49.518229Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:39:49.518229Z digest=sha256:3fe22451c456b3766c3af33c36356aabe4d2dc943d8239fbbe4eb14c220a2c9b