Pith. sign in

Paper Citation Record · LEDGER

Watching the AI Watchdogs: A Fairness and Robustness Analysis of AI Safety Moderation Classifiers

As of 11 August 2026, this Paper Citation Record lists 46 of 46 outbound references and 1 inbound Pith citation observation for arXiv:2501.13302.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2501.13302 v1

Coverage vector

measured 46 of 46 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-10T16:19:48.082702Z

measured 47 of 47 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-11T06:34:44.6726+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-10T18:41:54.081628Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-11T00:05:50.801220Z

Reference resolution

46 of 46 outbound references displayed

  • verified exact0
  • verified fuzzy4
  • unresolved42
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 957accdd-a6af-408f-9936-4feae9a131a4 · outbound

This paper cites GPT-4 Technical Report.

Watching the AI Watchdogs: A Fairness and Robustness Analysis of AI Safety Moderation Classifiers GPT-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-10T16:19:47.910889Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T16:19:47.910889Z digest=sha256:b645c197ac23bc2532966e82a03c7e0d868337317bd7407b97c3fb7159323e8f

Observation 994b7894-04d6-4678-b01f-86d99dcb43e4 · outbound

This paper cites PaLM 2 Technical Report.

Watching the AI Watchdogs: A Fairness and Robustness Analysis of AI Safety Moderation Classifiers PaLM 2 Technical Report

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-10T16:19:47.915189Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T16:19:47.915189Z digest=sha256:57f05c888441054de23279e407bf85b38490aa72eddb041455488361bd0c5a7b

Observation fdf210e6-af58-4c6f-9439-bbbd01208ec2 · outbound

This paper cites an unresolved cited work.

Watching the AI Watchdogs: A Fairness and Robustness Analysis of AI Safety Moderation Classifiers Unresolved cited work

Reference 3

Resolution
unresolved
raw_fallback, observed 2026-08-10T16:19:48.555144Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T16:19:47.919481Z digest=sha256:0727b7b5c86d4672fd17da8629956887e16af3a03b96392698ae44c4c8538eed

Observation b387d232-52ee-40ff-98bc-53a7a506ca64 · outbound

This paper cites an unresolved cited work.

Watching the AI Watchdogs: A Fairness and Robustness Analysis of AI Safety Moderation Classifiers Unresolved cited work

Reference 4

Resolution
unresolved
raw_fallback, observed 2026-08-10T16:19:48.543327Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T16:19:47.923447Z digest=sha256:28ec64e41c928b5af84981c2b5ad4d100a843fe14d6e7c7a21aad5cb5b5b4130

Observation 94bda893-9bd1-4600-a460-f0db83ae7d79 · outbound

This paper cites Measuring Political Bias in Large Language Models: What Is Said and How It Is Said.

Watching the AI Watchdogs: A Fairness and Robustness Analysis of AI Safety Moderation Classifiers Measuring Political Bias in Large Language Models: What Is Said and How It Is Said

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-10T16:19:47.927230Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T16:19:47.927230Z digest=sha256:913b0bf1047c2f2c62cba5ea7b4bb5852cd076a3f046e3ce8cb14861dff8bad6

Observation 6f5ed728-a3d1-4782-85af-0560f4be0bde · outbound

This paper cites an unresolved cited work.

Watching the AI Watchdogs: A Fairness and Robustness Analysis of AI Safety Moderation Classifiers Unresolved cited work

Reference 6

Resolution
unresolved
raw_fallback, observed 2026-08-10T16:19:48.531178Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T16:19:47.931343Z digest=sha256:cc37430112b7b5041397243f8c43b6780b2328651287258f071898209f8762a8

Observation 6aaee4bf-019f-4f79-b7d8-6cd9bc434e03 · outbound

This paper cites Machine Learning Robustness: A Primer.

Watching the AI Watchdogs: A Fairness and Robustness Analysis of AI Safety Moderation Classifiers Machine Learning Robustness: A Primer

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-10T16:19:47.935427Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T16:19:47.935427Z digest=sha256:c1d68e00c1f11575a56f1a7d6f757d0bf46b2077c280952bbbed43f8a9126a7a

Observation 40ba0447-1e27-4fa7-986e-7cfc3f37fe33 · outbound

This paper cites an unresolved cited work.

Watching the AI Watchdogs: A Fairness and Robustness Analysis of AI Safety Moderation Classifiers Unresolved cited work

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-10T16:19:47.939518Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T16:19:47.939518Z digest=sha256:46128a7ed8f85c14f67b828f3c3854a538ec4b9b6ec54f528ab51b46b4c329fa

Observation b9a4ae2d-2986-44c0-86ae-2c1bfcfde9ab · outbound

This paper cites an unresolved cited work.

Watching the AI Watchdogs: A Fairness and Robustness Analysis of AI Safety Moderation Classifiers Unresolved cited work

Reference 9

Resolution
unresolved
raw_fallback, observed 2026-08-10T16:19:48.511482Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T16:19:47.942816Z digest=sha256:b66a5be6be48adeb71f918110d7c44d7b5dd8f01e4fed0b85ae623da05915f58

Observation 5f466b0d-35e0-453c-9b36-1ab2527b282c · outbound

This paper cites an unresolved cited work.

Watching the AI Watchdogs: A Fairness and Robustness Analysis of AI Safety Moderation Classifiers Unresolved cited work

Reference 10

Resolution
unresolved
raw_fallback, observed 2026-08-10T16:19:48.499247Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T16:19:47.946449Z digest=sha256:2db26b1283d714d68020bed2c6caaf82742e13c7f05e6a0a77e4a7ea1e98b983

Observation 6d38a130-b64e-4e27-bce1-906b147fc153 · outbound

This paper cites an unresolved cited work.

Watching the AI Watchdogs: A Fairness and Robustness Analysis of AI Safety Moderation Classifiers Unresolved cited work

Reference 11

Resolution
unresolved
raw_fallback, observed 2026-08-10T16:19:48.486852Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T16:19:47.949889Z digest=sha256:83388f9b142723d3df031669fb3c7a103131ba90207a430774caafdadc917198

Observation 6d49dfe9-edeb-466e-88ca-990fdb6f3091 · outbound

This paper cites what data benefits my classifier?.

Watching the AI Watchdogs: A Fairness and Robustness Analysis of AI Safety Moderation Classifiers what data benefits my classifier?

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T16:19:48.475115Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T16:19:47.953417Z digest=sha256:5520608c017b857258dd99dd609eeb71e37d015130dd30fb4beb9c611a6639d5

Observation 01d6af94-8c05-4917-8193-7f0a93522893 · outbound

This paper cites an unresolved cited work.

Watching the AI Watchdogs: A Fairness and Robustness Analysis of AI Safety Moderation Classifiers Unresolved cited work

Reference 13

Resolution
unresolved
raw_fallback, observed 2026-08-10T16:19:48.463442Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T16:19:47.956787Z digest=sha256:bf40ccb128cd7d1b23530f2c52d95f6f584dde5fe01b5e4f17cf9e9bff6c1b63

Observation 0b158943-e49e-4194-a188-d673ef496a69 · outbound

This paper cites Towards F air V ideo S ummarization.

Watching the AI Watchdogs: A Fairness and Robustness Analysis of AI Safety Moderation Classifiers Towards F air V ideo S ummarization

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T16:19:48.451843Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T16:19:47.960199Z digest=sha256:8eaaa20e9090f8e8c238b79948aa313473c8f0ebbdc87f6cf9e1d3287c65c6fd

Observation 469152dd-23c9-4741-bc33-b57b8f86194a · outbound

This paper cites https://clarifai.com/clarifai/main/models/moderation-english-text-classification.

Watching the AI Watchdogs: A Fairness and Robustness Analysis of AI Safety Moderation Classifiers https://clarifai.com/clarifai/main/models/moderation-english-text-classification

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T16:19:48.439695Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T16:19:47.963889Z digest=sha256:50c7cfcb84e3dc3e6bcabd42bbd4114206e25b1b5244bc99a100f555ca6b4c75

Observation 833e1be4-6313-4157-b8b2-71756fda40d8 · outbound

This paper cites an unresolved cited work.

Watching the AI Watchdogs: A Fairness and Robustness Analysis of AI Safety Moderation Classifiers Unresolved cited work

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-10T16:19:47.967288Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T16:19:47.967288Z digest=sha256:2fe0d4dfa3faff451fd06052403dc8d9528c0b4a269516432175fd81df1ddad8

Observation 7eb2e53b-6f52-4a39-9029-3be0d4989878 · outbound

This paper cites an unresolved cited work.

Watching the AI Watchdogs: A Fairness and Robustness Analysis of AI Safety Moderation Classifiers Unresolved cited work

Reference 17

Resolution
unresolved
raw_fallback, observed 2026-08-10T16:19:48.420477Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T16:19:47.970743Z digest=sha256:52891a7f48591a8f0bf6dee9d6e95c764807e6f159d834fd5cd2bd694f834638

Observation 37171bea-4513-4777-a01f-b934a5d8eb45 · outbound

This paper cites an unresolved cited work.

Watching the AI Watchdogs: A Fairness and Robustness Analysis of AI Safety Moderation Classifiers Unresolved cited work

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-10T16:19:47.974101Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T16:19:47.974101Z digest=sha256:4d1b77fa223562c8a3accc3caf442be1e97a5a254a87849e38b2fed5867dcf01

Observation 4e01e04b-652c-40d9-b9fe-89ec5d468370 · outbound

This paper cites an unresolved cited work.

Watching the AI Watchdogs: A Fairness and Robustness Analysis of AI Safety Moderation Classifiers Unresolved cited work

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-10T16:19:47.977419Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T16:19:47.977419Z digest=sha256:e972ac2b95b165ea93a3074811c8f0aad6515c6ae84220ce645fed66b0af3e42

Observation fd54c1ef-1759-49df-99fb-9369e50a2a02 · outbound

This paper cites Building Guardrails for Large Language Models.

Watching the AI Watchdogs: A Fairness and Robustness Analysis of AI Safety Moderation Classifiers Building Guardrails for Large Language Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-10T16:19:47.980860Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T16:19:47.980860Z digest=sha256:30d2d5a843d17bfc7cb6dbe8d7cf1b984482fef0005f3218b6f353c0ffb35738

Observation 1c56b1aa-31a9-4a15-ab87-aa275dbc12f6 · outbound

This paper cites Towards Measuring the Representation of Subjective Global Opinions in Language Models.

Watching the AI Watchdogs: A Fairness and Robustness Analysis of AI Safety Moderation Classifiers Towards Measuring the Representation of Subjective Global Opinions in Language Models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-10T16:19:47.984874Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T16:19:47.984874Z digest=sha256:abef9ece5df84b5b042559ddb7f24326b846d29b5cbeedb741934854375f84b1

Observation ffe43d74-258e-45a0-ba09-5ebafb5619d2 · outbound

This paper cites an unresolved cited work.

Watching the AI Watchdogs: A Fairness and Robustness Analysis of AI Safety Moderation Classifiers Unresolved cited work

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-10T16:19:47.988841Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T16:19:47.988841Z digest=sha256:8d98ba1bfee1299701ed92191c77ae85360dc9f88bd167d43638a8a89d2c5588

Observation 40265bf1-2d53-4cea-b645-f9dc54041843 · outbound

This paper cites RealToxicityPrompts: Evaluating Neural Toxic Degeneration in Language Models.

Watching the AI Watchdogs: A Fairness and Robustness Analysis of AI Safety Moderation Classifiers RealToxicityPrompts: Evaluating Neural Toxic Degeneration in Language Models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-10T16:19:47.992480Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T16:19:47.992480Z digest=sha256:57db7b302c2c00953701004528c31c3a663590417d94bb77b748312475e906fc

Observation 4f7fc283-4ac4-4a98-8f4a-69a3027413e7 · outbound

This paper cites an unresolved cited work.

Watching the AI Watchdogs: A Fairness and Robustness Analysis of AI Safety Moderation Classifiers Unresolved cited work

Reference 24

Resolution
unresolved
raw_fallback, observed 2026-08-10T16:19:48.389553Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T16:19:47.996348Z digest=sha256:373cd848fabd0b86a2b3d240650bdc1a3123957fb986d81136375148c52f7e80

Observation 3120bf2e-c8d9-4288-96e8-ce3b560ca2d2 · outbound

This paper cites an unresolved cited work.

Watching the AI Watchdogs: A Fairness and Robustness Analysis of AI Safety Moderation Classifiers Unresolved cited work

Reference 25

Resolution
unresolved
raw_fallback, observed 2026-08-10T16:19:48.378264Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T16:19:47.999917Z digest=sha256:7ae95e560cee53139714f6a128b63528c15ba8eff5321814e206a264775094f6

Observation b5a3957c-06e0-43b3-a3ec-8e97510de234 · outbound

This paper cites BERTopic: Neural topic modeling with a class-based TF-IDF procedure.

Watching the AI Watchdogs: A Fairness and Robustness Analysis of AI Safety Moderation Classifiers BERTopic: Neural topic modeling with a class-based TF-IDF procedure

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-10T16:19:48.003786Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T16:19:48.003786Z digest=sha256:4017c30efef61a52fae719bc43bbb1d38d8e0d3cfb1b29a538ce674b14f5bc6e

Observation c3e0b635-8d6d-444f-89b2-965b6abd3673 · outbound

This paper cites an unresolved cited work.

Watching the AI Watchdogs: A Fairness and Robustness Analysis of AI Safety Moderation Classifiers Unresolved cited work

Reference 27

Resolution
unresolved
raw_fallback, observed 2026-08-10T16:19:48.367010Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T16:19:48.007666Z digest=sha256:7fcb55c1d62767a20064ad5483f0470ef58986f1181fde879cd1be176eb40a22

Observation d24dd804-9b07-43ef-b486-3422c813e1bc · outbound

This paper cites an unresolved cited work.

Watching the AI Watchdogs: A Fairness and Robustness Analysis of AI Safety Moderation Classifiers Unresolved cited work

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-10T16:19:48.011118Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T16:19:48.011118Z digest=sha256:ebd31800053b80debe5e65973d02aac757a1d35295d7856aae6368b80060c613

Observation 9f120699-f8b7-414f-992a-24852e974628 · outbound

This paper cites an unresolved cited work.

Watching the AI Watchdogs: A Fairness and Robustness Analysis of AI Safety Moderation Classifiers Unresolved cited work

Reference 29

Resolution
unresolved
raw_fallback, observed 2026-08-10T16:19:48.349006Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T16:19:48.014641Z digest=sha256:9d2ee1723b78cb62419f8dcce40d6c66b85d909e99bbac78572ede9e61b9877f

Observation 3b0804cd-4ab9-4b74-986a-422a78b03bca · outbound

This paper cites an unresolved cited work.

Watching the AI Watchdogs: A Fairness and Robustness Analysis of AI Safety Moderation Classifiers Unresolved cited work

Reference 30

Resolution
unresolved
raw_fallback, observed 2026-08-10T16:19:48.337606Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T16:19:48.018074Z digest=sha256:33df7fdad455477d8ff4280a3f8b55ac9efcbb5377ef354bf24681bdf00ec20b

Observation 735f0aa3-ae13-4d38-a391-cdb011a368c0 · outbound

This paper cites AutoDAN: Generating Stealthy Jailbreak Prompts on Aligned Large Language Models.

Watching the AI Watchdogs: A Fairness and Robustness Analysis of AI Safety Moderation Classifiers AutoDAN: Generating Stealthy Jailbreak Prompts on Aligned Large Language Models

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-10T16:19:48.024738Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T16:19:48.024738Z digest=sha256:8d581bd30e6561f87614f6fcc25b861c7e7afa7e617cefa4bd547d85259d90e4

Observation e62d96b7-2b9c-4e4c-a03b-dbf15871af17 · outbound

This paper cites An Empirical Study of Catastrophic Forgetting in Large Language Models During Continual Fine-tuning.

Watching the AI Watchdogs: A Fairness and Robustness Analysis of AI Safety Moderation Classifiers An Empirical Study of Catastrophic Forgetting in Large Language Models During Continual Fine-tuning

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-10T16:19:48.028776Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T16:19:48.028776Z digest=sha256:b9c5a6419eb9baa8229c69bea86ff4fbc2f65a2881298130ac306b2a94d6cadc

Observation f3ca1292-2bac-4aa4-91bc-0f9598203d2e · outbound

This paper cites an unresolved cited work.

Watching the AI Watchdogs: A Fairness and Robustness Analysis of AI Safety Moderation Classifiers Unresolved cited work

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-10T16:19:48.032831Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T16:19:48.032831Z digest=sha256:6c479dfc5c14d7992d4f13ee43911303ddaac0ad16a545a13ea56e6c87e49864

Observation 4c83c018-6e34-4c78-8582-53e0e3149d44 · outbound

This paper cites an unresolved cited work.

Watching the AI Watchdogs: A Fairness and Robustness Analysis of AI Safety Moderation Classifiers Unresolved cited work

Reference 34

Resolution
unresolved
raw_fallback, observed 2026-08-10T16:19:48.320838Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T16:19:48.036579Z digest=sha256:2fcd32089a5b8b9b7228df7f79d3745cc970c93bc9e34e2f623c9b52f8d259f3

Observation cc62cd94-0d3a-422c-b00f-eebbb86aed88 · outbound

This paper cites Toxic Bias: Perspective API Misreads German as More Toxic.

Watching the AI Watchdogs: A Fairness and Robustness Analysis of AI Safety Moderation Classifiers Toxic Bias: Perspective API Misreads German as More Toxic

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-10T16:19:48.040346Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T16:19:48.040346Z digest=sha256:627bba39344514b8561a70b4d7bb7c3203df05a42052ad7b076d0fbb438346ec

Observation 7957712b-2d95-451e-83a7-b446cda18662 · outbound

This paper cites https://platform.openai.com/docs/guides/moderation.

Watching the AI Watchdogs: A Fairness and Robustness Analysis of AI Safety Moderation Classifiers https://platform.openai.com/docs/guides/moderation

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T16:19:48.309591Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T16:19:48.044231Z digest=sha256:4a8d43eeeceaefc99360e15efe6f0b3bcfc63e255db29c8e30eef1a20ff722a1

Observation c8a19fc6-0376-434a-b9a3-6de65f6886d0 · outbound

This paper cites an unresolved cited work.

Watching the AI Watchdogs: A Fairness and Robustness Analysis of AI Safety Moderation Classifiers Unresolved cited work

Reference 37

Resolution
unresolved
raw_fallback, observed 2026-08-10T16:19:48.298317Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T16:19:48.048144Z digest=sha256:6209853f9cd0539859fdc223181534bbaa2d327528ac1a4926c39e3437e7a996

Observation 81c6a5ad-2387-4b42-af6e-661b23d00fd4 · outbound

This paper cites Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!.

Watching the AI Watchdogs: A Fairness and Robustness Analysis of AI Safety Moderation Classifiers Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-10T16:19:48.051712Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T16:19:48.051712Z digest=sha256:bc93e0bff42f8f70c4621d1a0e0b88f50feecd00fd31d5ba216f609c723e4b08

Observation 87807929-10c4-4f5a-acdf-3ba5921ceb0e · outbound

This paper cites an unresolved cited work.

Watching the AI Watchdogs: A Fairness and Robustness Analysis of AI Safety Moderation Classifiers Unresolved cited work

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-10T16:19:48.055948Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T16:19:48.055948Z digest=sha256:7f0e97c67b6713c7f8358676a6fe63e30c5089bdc6eff6d38ab5e2d2ca63ddad

Observation be32b94c-a11a-42f4-91ad-6f8b2e6f96a6 · outbound

This paper cites The Woman Worked as a Babysitter: On Biases in Language Generation.

Watching the AI Watchdogs: A Fairness and Robustness Analysis of AI Safety Moderation Classifiers The Woman Worked as a Babysitter: On Biases in Language Generation

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-10T16:19:48.059776Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T16:19:48.059776Z digest=sha256:8653e165a9a413f95469e57443fe6a153b72feb0b8f2e342f52782738c9a84bb

Observation 5501aa05-f5be-4440-b8c3-fc48a7402fea · outbound

This paper cites an unresolved cited work.

Watching the AI Watchdogs: A Fairness and Robustness Analysis of AI Safety Moderation Classifiers Unresolved cited work

Reference 41

Resolution
unresolved
raw_fallback, observed 2026-08-10T16:19:48.287544Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T16:19:48.063813Z digest=sha256:22f8204f44323bfe3d4381e7b8018b6ca3d06b5b1f0685db108a5080c3bd9b27

Observation 8308824a-6f8c-404f-a3d3-d3f10e695b8d · outbound

This paper cites A Comprehensive Capability Analysis of GPT-3 and GPT-3.5 Series Models.

Watching the AI Watchdogs: A Fairness and Robustness Analysis of AI Safety Moderation Classifiers A Comprehensive Capability Analysis of GPT-3 and GPT-3.5 Series Models

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-10T16:19:48.067303Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T16:19:48.067303Z digest=sha256:e7518261db1fcdbb7ebb155316bcc839793b21c095ed35defd60127f6ecc7a44

Observation 1c7a6349-470a-42d5-a062-b6dd14613b53 · outbound

This paper cites How Johnny Can Persuade LLMs to Jailbreak Them: Rethinking Persuasion to Challenge AI Safety by Humanizing LLMs.

Watching the AI Watchdogs: A Fairness and Robustness Analysis of AI Safety Moderation Classifiers How Johnny Can Persuade LLMs to Jailbreak Them: Rethinking Persuasion to Challenge AI Safety by Humanizing LLMs

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-10T16:19:48.071026Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T16:19:48.071026Z digest=sha256:77a95f5f091fad10fc0b2a64676a34f20d8cfd1c1a42f68a2e70a793e5808315

Observation 232e39aa-c4af-45bb-a4fa-bde25f02c480 · outbound

This paper cites Universal and Transferable Adversarial Attacks on Aligned Language Models.

Watching the AI Watchdogs: A Fairness and Robustness Analysis of AI Safety Moderation Classifiers Universal and Transferable Adversarial Attacks on Aligned Language Models

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-10T16:19:48.074807Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T16:19:48.074807Z digest=sha256:3cb967307be5010a6b0e94c0cd1ca76b865fdf0befde7e6a13994b2c8a327b87

Observation 4a1c3282-f29a-4686-a57d-3e52fbfca083 · outbound

This paper cites URL: " 'urlintro :=.

Watching the AI Watchdogs: A Fairness and Robustness Analysis of AI Safety Moderation Classifiers URL: " 'urlintro :=

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-10T16:19:48.078510Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T16:19:48.078510Z digest=sha256:250153a4fd0fe9952877213fb3919bad4242858bf30aae4a674586e784353c68

Observation 8d0b8269-2142-4f6b-92f0-ff69ef57c879 · outbound

This paper cites write newline.

Watching the AI Watchdogs: A Fairness and Robustness Analysis of AI Safety Moderation Classifiers write newline

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-10T16:19:48.082702Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T16:19:48.082702Z digest=sha256:d1103044a9aeda00971af37fdfc607dbf2a51b33f4935d9cc70ca3528995e2b1

Pith citing papers

Observation cae08519-4457-440f-886f-b526fbb18938 · inbound

To Lie or Not to Lie? Investigating The Biased Spread of Global Lies by LLMs cites this paper.

To Lie or Not to Lie? Investigating The Biased Spread of Global Lies by LLMs Watching the AI Watchdogs: A Fairness and Robustness Analysis of AI Safety Moderation Classifiers

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-11T00:05:50.803437Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-05-10T18:41:54.081628Z digest=sha256:31589aa9799c411e16b222bfb042d1e780426a2955885ca5860b2efd8604325b