Pith. sign in

Paper Citation Record · LEDGER

Watching the AI Watchdogs: A Fairness and Robustness Analysis of AI Safety Moderation Classifiers

As of 11 August 2026, this Paper Citation Record lists 46 of 46 outbound references and 1 inbound Pith citation observation for arXiv:2501.13302.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2501.13302 v1

Coverage vector

measured 46 of 46 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-10T16:19:48.082702Z

measured 47 of 47 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-11T06:34:44.6726+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-10T18:41:54.081628Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-11T00:05:50.801220Z

Reference resolution

46 of 46 outbound references displayed

  • verified exact0
  • verified fuzzy4
  • unresolved42
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 957accdd-a6af-408f-9936-4feae9a131a4 · outbound

This paper cites GPT-4 Technical Report.

Watching the AI Watchdogs: A Fairness and Robustness Analysis of AI Safety Moderation Classifiers GPT-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-10T16:19:47.910889Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T16:19:47.910889Z digest=sha256:5dade66be5c356f902171a26b53e4cb65abb71e7253febb07fe2dddeb5017eee

Observation 994b7894-04d6-4678-b01f-86d99dcb43e4 · outbound

This paper cites PaLM 2 Technical Report.

Watching the AI Watchdogs: A Fairness and Robustness Analysis of AI Safety Moderation Classifiers PaLM 2 Technical Report

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-10T16:19:47.915189Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T16:19:47.915189Z digest=sha256:96ef8704cb577d26a9d9b6a6a3224dc4e53122cae720cb84d95e3b43113b83f6

Observation fdf210e6-af58-4c6f-9439-bbbd01208ec2 · outbound

This paper cites an unresolved cited work.

Watching the AI Watchdogs: A Fairness and Robustness Analysis of AI Safety Moderation Classifiers Unresolved cited work

Reference 3

Resolution
unresolved
raw_fallback, observed 2026-08-10T16:19:48.555144Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T16:19:47.919481Z digest=sha256:a4d9ecd9651b7a5575c95bc08aa24df2eee538843ab127e5f12bc98e9094a9bd

Observation b387d232-52ee-40ff-98bc-53a7a506ca64 · outbound

This paper cites an unresolved cited work.

Watching the AI Watchdogs: A Fairness and Robustness Analysis of AI Safety Moderation Classifiers Unresolved cited work

Reference 4

Resolution
unresolved
raw_fallback, observed 2026-08-10T16:19:48.543327Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T16:19:47.923447Z digest=sha256:434ee2109de04362ef53c91c3008b666607cb259fe96160c2549cc07ee9f300f

Observation 94bda893-9bd1-4600-a460-f0db83ae7d79 · outbound

This paper cites Measuring Political Bias in Large Language Models: What Is Said and How It Is Said.

Watching the AI Watchdogs: A Fairness and Robustness Analysis of AI Safety Moderation Classifiers Measuring Political Bias in Large Language Models: What Is Said and How It Is Said

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-10T16:19:47.927230Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T16:19:47.927230Z digest=sha256:09c09ac6761883c6153c645843b1605f8e9a8fa3ed237b0f66f73ae525e69cb7

Observation 6f5ed728-a3d1-4782-85af-0560f4be0bde · outbound

This paper cites an unresolved cited work.

Watching the AI Watchdogs: A Fairness and Robustness Analysis of AI Safety Moderation Classifiers Unresolved cited work

Reference 6

Resolution
unresolved
raw_fallback, observed 2026-08-10T16:19:48.531178Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T16:19:47.931343Z digest=sha256:c132345ab17fb879d810be969d38fd7d1812c74189c58ac41685ce0725c64d7b

Observation 6aaee4bf-019f-4f79-b7d8-6cd9bc434e03 · outbound

This paper cites Machine Learning Robustness: A Primer.

Watching the AI Watchdogs: A Fairness and Robustness Analysis of AI Safety Moderation Classifiers Machine Learning Robustness: A Primer

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-10T16:19:47.935427Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T16:19:47.935427Z digest=sha256:b8c6b02095f0cd2339399df4ae00d46f87b60f204ca40d5e0514c97a26aa9436

Observation 40ba0447-1e27-4fa7-986e-7cfc3f37fe33 · outbound

This paper cites an unresolved cited work.

Watching the AI Watchdogs: A Fairness and Robustness Analysis of AI Safety Moderation Classifiers Unresolved cited work

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-10T16:19:47.939518Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T16:19:47.939518Z digest=sha256:b509f7739a688b2983f9abb7ec5342338193a69886f676d901634ef5acd24045

Observation b9a4ae2d-2986-44c0-86ae-2c1bfcfde9ab · outbound

This paper cites an unresolved cited work.

Watching the AI Watchdogs: A Fairness and Robustness Analysis of AI Safety Moderation Classifiers Unresolved cited work

Reference 9

Resolution
unresolved
raw_fallback, observed 2026-08-10T16:19:48.511482Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T16:19:47.942816Z digest=sha256:b4d8a3e04b4cb5c6846515aa8182185cd945feda9e6e94f98b460f17b87ba885

Observation 5f466b0d-35e0-453c-9b36-1ab2527b282c · outbound

This paper cites an unresolved cited work.

Watching the AI Watchdogs: A Fairness and Robustness Analysis of AI Safety Moderation Classifiers Unresolved cited work

Reference 10

Resolution
unresolved
raw_fallback, observed 2026-08-10T16:19:48.499247Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T16:19:47.946449Z digest=sha256:6552d2585be42410067bc79e175d80e4a406a362f7324b1f3fb96b666d325f5e

Observation 6d38a130-b64e-4e27-bce1-906b147fc153 · outbound

This paper cites an unresolved cited work.

Watching the AI Watchdogs: A Fairness and Robustness Analysis of AI Safety Moderation Classifiers Unresolved cited work

Reference 11

Resolution
unresolved
raw_fallback, observed 2026-08-10T16:19:48.486852Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T16:19:47.949889Z digest=sha256:e733b12ade177671446209bbb04ec6fb32d3e96001002760daf63dbf74fc4637

Observation 6d49dfe9-edeb-466e-88ca-990fdb6f3091 · outbound

This paper cites what data benefits my classifier?.

Watching the AI Watchdogs: A Fairness and Robustness Analysis of AI Safety Moderation Classifiers what data benefits my classifier?

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T16:19:48.475115Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T16:19:47.953417Z digest=sha256:86087ca101aea7d36318b80a60e44492adf6379496f0b90911a6514ee538cda1

Observation 01d6af94-8c05-4917-8193-7f0a93522893 · outbound

This paper cites an unresolved cited work.

Watching the AI Watchdogs: A Fairness and Robustness Analysis of AI Safety Moderation Classifiers Unresolved cited work

Reference 13

Resolution
unresolved
raw_fallback, observed 2026-08-10T16:19:48.463442Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T16:19:47.956787Z digest=sha256:eb763ea61a73982b4f9662a2a15977dd1c7c81459b3c031fb96c019da7e2a2f1

Observation 0b158943-e49e-4194-a188-d673ef496a69 · outbound

This paper cites Towards F air V ideo S ummarization.

Watching the AI Watchdogs: A Fairness and Robustness Analysis of AI Safety Moderation Classifiers Towards F air V ideo S ummarization

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T16:19:48.451843Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T16:19:47.960199Z digest=sha256:546fd25e90a1664a81858c3bf126e4caf36fe49388c5476c597e3461ed294fd6

Observation 469152dd-23c9-4741-bc33-b57b8f86194a · outbound

This paper cites https://clarifai.com/clarifai/main/models/moderation-english-text-classification.

Watching the AI Watchdogs: A Fairness and Robustness Analysis of AI Safety Moderation Classifiers https://clarifai.com/clarifai/main/models/moderation-english-text-classification

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T16:19:48.439695Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T16:19:47.963889Z digest=sha256:e5f30baac45f00d6ca10d36db9131f461ebd7d4ce0d8c9a636e8e1a12d657940

Observation 833e1be4-6313-4157-b8b2-71756fda40d8 · outbound

This paper cites an unresolved cited work.

Watching the AI Watchdogs: A Fairness and Robustness Analysis of AI Safety Moderation Classifiers Unresolved cited work

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-10T16:19:47.967288Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T16:19:47.967288Z digest=sha256:c3481015c23b2f9bf4b0f86b002acbc3dcc9aa05efc047dfcccd8eb942217ead

Observation 7eb2e53b-6f52-4a39-9029-3be0d4989878 · outbound

This paper cites an unresolved cited work.

Watching the AI Watchdogs: A Fairness and Robustness Analysis of AI Safety Moderation Classifiers Unresolved cited work

Reference 17

Resolution
unresolved
raw_fallback, observed 2026-08-10T16:19:48.420477Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T16:19:47.970743Z digest=sha256:fe7f4c1fa86759bac8a7dad178f9daafeccc5006ebb5613cb7faa13ecd2bb303

Observation 37171bea-4513-4777-a01f-b934a5d8eb45 · outbound

This paper cites an unresolved cited work.

Watching the AI Watchdogs: A Fairness and Robustness Analysis of AI Safety Moderation Classifiers Unresolved cited work

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-10T16:19:47.974101Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T16:19:47.974101Z digest=sha256:762496ca9d714fdc371a68f9a91e2226640651e84844dde7a7bdc15e0ab1966a

Observation 4e01e04b-652c-40d9-b9fe-89ec5d468370 · outbound

This paper cites an unresolved cited work.

Watching the AI Watchdogs: A Fairness and Robustness Analysis of AI Safety Moderation Classifiers Unresolved cited work

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-10T16:19:47.977419Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T16:19:47.977419Z digest=sha256:a72f7a6b059aefbe62a64a25fe201d82d706021b2a53a5b2cb2fd358a0482d25

Observation fd54c1ef-1759-49df-99fb-9369e50a2a02 · outbound

This paper cites Building Guardrails for Large Language Models.

Watching the AI Watchdogs: A Fairness and Robustness Analysis of AI Safety Moderation Classifiers Building Guardrails for Large Language Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-10T16:19:47.980860Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T16:19:47.980860Z digest=sha256:d8aab48502a7057f1647bb99843b117bdb5f1810efc05044a808b85249d8e9f8

Observation 1c56b1aa-31a9-4a15-ab87-aa275dbc12f6 · outbound

This paper cites Towards Measuring the Representation of Subjective Global Opinions in Language Models.

Watching the AI Watchdogs: A Fairness and Robustness Analysis of AI Safety Moderation Classifiers Towards Measuring the Representation of Subjective Global Opinions in Language Models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-10T16:19:47.984874Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T16:19:47.984874Z digest=sha256:fb84fcd3d0a41f7d6003f1ce6c2ab71c85a3fcde9531008a85212ec50513c64a

Observation ffe43d74-258e-45a0-ba09-5ebafb5619d2 · outbound

This paper cites an unresolved cited work.

Watching the AI Watchdogs: A Fairness and Robustness Analysis of AI Safety Moderation Classifiers Unresolved cited work

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-10T16:19:47.988841Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T16:19:47.988841Z digest=sha256:d1ecf7b9dee106b719e3b93c7034c8948bb4c3f25b8482f2b3cddfb0aea253e5

Observation 40265bf1-2d53-4cea-b645-f9dc54041843 · outbound

This paper cites RealToxicityPrompts: Evaluating Neural Toxic Degeneration in Language Models.

Watching the AI Watchdogs: A Fairness and Robustness Analysis of AI Safety Moderation Classifiers RealToxicityPrompts: Evaluating Neural Toxic Degeneration in Language Models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-10T16:19:47.992480Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T16:19:47.992480Z digest=sha256:8e0dc3d7c3c31841a06d182a7b35170c9a7c47dc7160944dfa668cf8ebd355e1

Observation 4f7fc283-4ac4-4a98-8f4a-69a3027413e7 · outbound

This paper cites an unresolved cited work.

Watching the AI Watchdogs: A Fairness and Robustness Analysis of AI Safety Moderation Classifiers Unresolved cited work

Reference 24

Resolution
unresolved
raw_fallback, observed 2026-08-10T16:19:48.389553Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T16:19:47.996348Z digest=sha256:48e07052be0585c93969fa0d48fc956d61231449f9405cd4edbf0888dafd6573

Observation 3120bf2e-c8d9-4288-96e8-ce3b560ca2d2 · outbound

This paper cites an unresolved cited work.

Watching the AI Watchdogs: A Fairness and Robustness Analysis of AI Safety Moderation Classifiers Unresolved cited work

Reference 25

Resolution
unresolved
raw_fallback, observed 2026-08-10T16:19:48.378264Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T16:19:47.999917Z digest=sha256:c695c9df29798f4b2957728d3a1633a63b8723a3668c2796911e209fde37ece3

Observation b5a3957c-06e0-43b3-a3ec-8e97510de234 · outbound

This paper cites BERTopic: Neural topic modeling with a class-based TF-IDF procedure.

Watching the AI Watchdogs: A Fairness and Robustness Analysis of AI Safety Moderation Classifiers BERTopic: Neural topic modeling with a class-based TF-IDF procedure

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-10T16:19:48.003786Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T16:19:48.003786Z digest=sha256:465937fcb8d51027b13335b56f93cedb87a8376a516ecd82ca4a46fe0da82ebf

Observation c3e0b635-8d6d-444f-89b2-965b6abd3673 · outbound

This paper cites an unresolved cited work.

Watching the AI Watchdogs: A Fairness and Robustness Analysis of AI Safety Moderation Classifiers Unresolved cited work

Reference 27

Resolution
unresolved
raw_fallback, observed 2026-08-10T16:19:48.367010Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T16:19:48.007666Z digest=sha256:4196a03158a95f68cc4e914fdb3fe11eae4bcdf58ae15bc9b3d0a9b6b83b05fb

Observation d24dd804-9b07-43ef-b486-3422c813e1bc · outbound

This paper cites an unresolved cited work.

Watching the AI Watchdogs: A Fairness and Robustness Analysis of AI Safety Moderation Classifiers Unresolved cited work

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-10T16:19:48.011118Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T16:19:48.011118Z digest=sha256:b5170137e5b8742a1e5f9a054f73f0862a2760c6b08327d9c542fb6a6877fb42

Observation 9f120699-f8b7-414f-992a-24852e974628 · outbound

This paper cites an unresolved cited work.

Watching the AI Watchdogs: A Fairness and Robustness Analysis of AI Safety Moderation Classifiers Unresolved cited work

Reference 29

Resolution
unresolved
raw_fallback, observed 2026-08-10T16:19:48.349006Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T16:19:48.014641Z digest=sha256:d1adfce6f9b558de4696324c342f88460e6bba96bc521110c3cc832f33bbfe9a

Observation 3b0804cd-4ab9-4b74-986a-422a78b03bca · outbound

This paper cites an unresolved cited work.

Watching the AI Watchdogs: A Fairness and Robustness Analysis of AI Safety Moderation Classifiers Unresolved cited work

Reference 30

Resolution
unresolved
raw_fallback, observed 2026-08-10T16:19:48.337606Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T16:19:48.018074Z digest=sha256:6618f911f0dcc9dc4a0481544d22e085e768e813128776a616058f9819da1def

Observation 735f0aa3-ae13-4d38-a391-cdb011a368c0 · outbound

This paper cites AutoDAN: Generating Stealthy Jailbreak Prompts on Aligned Large Language Models.

Watching the AI Watchdogs: A Fairness and Robustness Analysis of AI Safety Moderation Classifiers AutoDAN: Generating Stealthy Jailbreak Prompts on Aligned Large Language Models

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-10T16:19:48.024738Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T16:19:48.024738Z digest=sha256:74aedd6ddb39d302739f008f8b48f9859d8d9553e9bb4b5b2a757aca865e682b

Observation e62d96b7-2b9c-4e4c-a03b-dbf15871af17 · outbound

This paper cites An Empirical Study of Catastrophic Forgetting in Large Language Models During Continual Fine-tuning.

Watching the AI Watchdogs: A Fairness and Robustness Analysis of AI Safety Moderation Classifiers An Empirical Study of Catastrophic Forgetting in Large Language Models During Continual Fine-tuning

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-10T16:19:48.028776Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T16:19:48.028776Z digest=sha256:9d837b3fe1571260196ee5b6fe06401688ff123f3063073f321d868b17af7a44

Observation f3ca1292-2bac-4aa4-91bc-0f9598203d2e · outbound

This paper cites an unresolved cited work.

Watching the AI Watchdogs: A Fairness and Robustness Analysis of AI Safety Moderation Classifiers Unresolved cited work

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-10T16:19:48.032831Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T16:19:48.032831Z digest=sha256:a8614b4f886f871b408880b60b0b5ec23441cafead1ea332a80814cd8b89a906

Observation 4c83c018-6e34-4c78-8582-53e0e3149d44 · outbound

This paper cites an unresolved cited work.

Watching the AI Watchdogs: A Fairness and Robustness Analysis of AI Safety Moderation Classifiers Unresolved cited work

Reference 34

Resolution
unresolved
raw_fallback, observed 2026-08-10T16:19:48.320838Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T16:19:48.036579Z digest=sha256:7c10351ff928826361d2a53543cf5a59d0cc0d3ab107875aea0d2d914c53e486

Observation cc62cd94-0d3a-422c-b00f-eebbb86aed88 · outbound

This paper cites Toxic Bias: Perspective API Misreads German as More Toxic.

Watching the AI Watchdogs: A Fairness and Robustness Analysis of AI Safety Moderation Classifiers Toxic Bias: Perspective API Misreads German as More Toxic

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-10T16:19:48.040346Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T16:19:48.040346Z digest=sha256:0c93e7161b696ed0a13ce01523b151b21620b0f4fd9d923aa15f42f52d32009e

Observation 7957712b-2d95-451e-83a7-b446cda18662 · outbound

This paper cites https://platform.openai.com/docs/guides/moderation.

Watching the AI Watchdogs: A Fairness and Robustness Analysis of AI Safety Moderation Classifiers https://platform.openai.com/docs/guides/moderation

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T16:19:48.309591Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T16:19:48.044231Z digest=sha256:7c091bb6a47bc6dacfeb392216630497c1f3930d0c1ec2d3c94999127807600b

Observation c8a19fc6-0376-434a-b9a3-6de65f6886d0 · outbound

This paper cites an unresolved cited work.

Watching the AI Watchdogs: A Fairness and Robustness Analysis of AI Safety Moderation Classifiers Unresolved cited work

Reference 37

Resolution
unresolved
raw_fallback, observed 2026-08-10T16:19:48.298317Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T16:19:48.048144Z digest=sha256:df88b9e7fd2f04ae3513e319d9b80f4beafcb8a9edc3454db832f7ab60269d41

Observation 81c6a5ad-2387-4b42-af6e-661b23d00fd4 · outbound

This paper cites Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!.

Watching the AI Watchdogs: A Fairness and Robustness Analysis of AI Safety Moderation Classifiers Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-10T16:19:48.051712Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T16:19:48.051712Z digest=sha256:d0da2b31870469e2242c493971e78e4f0875096eaa806d8a80f5173145be5f0c

Observation 87807929-10c4-4f5a-acdf-3ba5921ceb0e · outbound

This paper cites an unresolved cited work.

Watching the AI Watchdogs: A Fairness and Robustness Analysis of AI Safety Moderation Classifiers Unresolved cited work

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-10T16:19:48.055948Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T16:19:48.055948Z digest=sha256:9ade7284189abfe6d599ea0c9f6eb4bb6a8453cd1337b5a0f5322debc2c2e385

Observation be32b94c-a11a-42f4-91ad-6f8b2e6f96a6 · outbound

This paper cites The Woman Worked as a Babysitter: On Biases in Language Generation.

Watching the AI Watchdogs: A Fairness and Robustness Analysis of AI Safety Moderation Classifiers The Woman Worked as a Babysitter: On Biases in Language Generation

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-10T16:19:48.059776Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T16:19:48.059776Z digest=sha256:e0aea6c4026d8a55be407df5fe7e676a46569d335a7bae08a3c2fd03def63305

Observation 5501aa05-f5be-4440-b8c3-fc48a7402fea · outbound

This paper cites an unresolved cited work.

Watching the AI Watchdogs: A Fairness and Robustness Analysis of AI Safety Moderation Classifiers Unresolved cited work

Reference 41

Resolution
unresolved
raw_fallback, observed 2026-08-10T16:19:48.287544Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T16:19:48.063813Z digest=sha256:ea0b6bea96a55758e09c5046744331e8aeb6aeb6d8d47cc4e61ce13af6448491

Observation 8308824a-6f8c-404f-a3d3-d3f10e695b8d · outbound

This paper cites A Comprehensive Capability Analysis of GPT-3 and GPT-3.5 Series Models.

Watching the AI Watchdogs: A Fairness and Robustness Analysis of AI Safety Moderation Classifiers A Comprehensive Capability Analysis of GPT-3 and GPT-3.5 Series Models

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-10T16:19:48.067303Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T16:19:48.067303Z digest=sha256:3916ae12d6d239092f2c69817202dad7bd88805af99b07f1e512bf720ede11bd

Observation 1c7a6349-470a-42d5-a062-b6dd14613b53 · outbound

This paper cites How Johnny Can Persuade LLMs to Jailbreak Them: Rethinking Persuasion to Challenge AI Safety by Humanizing LLMs.

Watching the AI Watchdogs: A Fairness and Robustness Analysis of AI Safety Moderation Classifiers How Johnny Can Persuade LLMs to Jailbreak Them: Rethinking Persuasion to Challenge AI Safety by Humanizing LLMs

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-10T16:19:48.071026Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T16:19:48.071026Z digest=sha256:51bc484896372586c973c829229053cd1a40b97aaf18d73bd997c80108df45df

Observation 232e39aa-c4af-45bb-a4fa-bde25f02c480 · outbound

This paper cites Universal and Transferable Adversarial Attacks on Aligned Language Models.

Watching the AI Watchdogs: A Fairness and Robustness Analysis of AI Safety Moderation Classifiers Universal and Transferable Adversarial Attacks on Aligned Language Models

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-10T16:19:48.074807Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T16:19:48.074807Z digest=sha256:bc8ac6d6e5244e2b45e08da3818774d142ad01eefbeea9c8cdc032b611cdaa32

Observation 4a1c3282-f29a-4686-a57d-3e52fbfca083 · outbound

This paper cites URL: " 'urlintro :=.

Watching the AI Watchdogs: A Fairness and Robustness Analysis of AI Safety Moderation Classifiers URL: " 'urlintro :=

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-10T16:19:48.078510Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T16:19:48.078510Z digest=sha256:7d44a09d680b9087cfab2bd9271aa624b97760b4dc66e6774fb5a6464178c75b

Observation 8d0b8269-2142-4f6b-92f0-ff69ef57c879 · outbound

This paper cites write newline.

Watching the AI Watchdogs: A Fairness and Robustness Analysis of AI Safety Moderation Classifiers write newline

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-10T16:19:48.082702Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T16:19:48.082702Z digest=sha256:1a0ff27713611187a9b18acef8f5dd21e37ca02b6218b3c2092ef93b3a2d8401

Pith citing papers

Observation cae08519-4457-440f-886f-b526fbb18938 · inbound

To Lie or Not to Lie? Investigating The Biased Spread of Global Lies by LLMs cites this paper.

To Lie or Not to Lie? Investigating The Biased Spread of Global Lies by LLMs Watching the AI Watchdogs: A Fairness and Robustness Analysis of AI Safety Moderation Classifiers

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-11T00:05:50.803437Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-05-10T18:41:54.081628Z digest=sha256:1de4d80c85e65282ae0a027717d858e1d91d132681938f6dcce73f56ae873a6c