Pith. sign in

Paper Citation Record · LEDGER

The Scales of Justitia: A Comprehensive Survey on Safety Evaluation of LLMs

As of 11 August 2026, this Paper Citation Record lists 100 of 215 outbound references and 9 inbound Pith citation observations for arXiv:2506.11094.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.11094 v2

Coverage vector

measured 100 of 215 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T10:17:26.896104Z

measured 109 of 109 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-11T06:34:44.6726+00:00

measured 9 of 9 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-08T18:47:38.262660Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-04T03:39:30.627528Z

Reference resolution

100 of 215 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved100
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 0d671c9f-27dc-4248-92ca-ef486691a2b3 · outbound

This paper cites A survey on evaluation of large language models,.

The Scales of Justitia: A Comprehensive Survey on Safety Evaluation of LLMs A survey on evaluation of large language models,

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T10:17:26.578627Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:17:26.578627Z digest=sha256:f92765dfaaac303cfabf56882852a9525ee9d2988d3d617165ed6d517fcaabdf

Observation 79aab667-973b-4619-8aae-7c2fb795b132 · outbound

This paper cites A systematic review on human and computer interaction,.

The Scales of Justitia: A Comprehensive Survey on Safety Evaluation of LLMs A systematic review on human and computer interaction,

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T10:17:26.583802Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:17:26.583802Z digest=sha256:a760972f143922cc6643b25f45e762f8a38655b17680f8f28fb5e507d39637e0

Observation 7dee63a2-a51e-4896-84cb-287a56989c97 · outbound

This paper cites From Handcrafted Features to LLMs: A Brief Survey for Machine Translation Quality Estimation.

The Scales of Justitia: A Comprehensive Survey on Safety Evaluation of LLMs From Handcrafted Features to LLMs: A Brief Survey for Machine Translation Quality Estimation

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T10:17:26.587464Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:17:26.587464Z digest=sha256:36c3e49fb744a9f06088dae38d58bdf451f560e3404ef056118cf7c58fb987a3

Observation 794a941e-1db0-4fa3-8c6f-cdc41b73ba3e · outbound

This paper cites A Survey on Evaluating Large Language Models in Code Generation Tasks.

The Scales of Justitia: A Comprehensive Survey on Safety Evaluation of LLMs A Survey on Evaluating Large Language Models in Code Generation Tasks

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T10:17:26.591222Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:17:26.591222Z digest=sha256:3185b7ae11ae88acfdca7181dacecd17814d71d0aa646cfa9d12e28bea82ebb1

Observation 8f2b867d-b91f-44e1-bce2-84835ad69aa5 · outbound

This paper cites Beyond Self-Talk: A Communication-Centric Survey of LLM-Based Multi-Agent Systems.

The Scales of Justitia: A Comprehensive Survey on Safety Evaluation of LLMs Beyond Self-Talk: A Communication-Centric Survey of LLM-Based Multi-Agent Systems

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T10:17:26.594818Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:17:26.594818Z digest=sha256:23db55b6927c8df4626c724d717174318a4fa5a5aba9da26507b8165cb1ab8f5

Observation e2cf32aa-d78e-41dc-a8cc-284c5133b36b · outbound

This paper cites Unleash LLMs Potential for Recommendation by Coordinating Twin-Tower Dynamic Semantic Token Generator.

The Scales of Justitia: A Comprehensive Survey on Safety Evaluation of LLMs Unleash LLMs Potential for Recommendation by Coordinating Twin-Tower Dynamic Semantic Token Generator

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T10:17:26.598372Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:17:26.598372Z digest=sha256:c86052bc3ec1f65c0260ed77ab6af526b21d52eeac8410b1c6cd882f568d3edd

Observation 3608d722-55aa-45b5-8579-e4e6c896ec9e · outbound

This paper cites Beyond Accuracy: Evaluating the Reasoning Behavior of Large Language Models -- A Survey.

The Scales of Justitia: A Comprehensive Survey on Safety Evaluation of LLMs Beyond Accuracy: Evaluating the Reasoning Behavior of Large Language Models -- A Survey

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T10:17:26.602021Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:17:26.602021Z digest=sha256:a506348bd2b46d7562c7732696b62db47699d8b1e1c3ebe30a80fc898bc030c6

Observation 3c8b545b-a87d-4e72-8e17-c060b15cf272 · outbound

This paper cites Universal and Transferable Adversarial Attacks on Aligned Language Models.

The Scales of Justitia: A Comprehensive Survey on Safety Evaluation of LLMs Universal and Transferable Adversarial Attacks on Aligned Language Models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T10:17:26.605878Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:17:26.605878Z digest=sha256:d8b0763e7b288b5aebf22788de6ccba4376e68685cc34f64309a3428dd36a07f

Observation 07958821-5743-4529-b082-170c73e16931 · outbound

This paper cites Is chatgpt fair for recommendation? evaluating fairness in large language model recommendation,.

The Scales of Justitia: A Comprehensive Survey on Safety Evaluation of LLMs Is chatgpt fair for recommendation? evaluating fairness in large language model recommendation,

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T10:17:26.609382Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:17:26.609382Z digest=sha256:51612d8f5f0d25fd0e68b02e2d191161ca9de848d8c1892b882ead7ab35f5fb0

Observation 9fc1dd70-6916-4333-b1ca-16117d61bcc0 · outbound

This paper cites FineFake: A Knowledge-Enriched Dataset for Fine-Grained Multi-Domain Fake News Detection.

The Scales of Justitia: A Comprehensive Survey on Safety Evaluation of LLMs FineFake: A Knowledge-Enriched Dataset for Fine-Grained Multi-Domain Fake News Detection

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T10:17:26.612552Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:17:26.612552Z digest=sha256:e2217c0d94bd8e0fbbbf535ea765071b7575e819559503d81af4325befcf0a67

Observation 4a85706f-d7d3-4a80-9b79-02a63d2b3316 · outbound

This paper cites Baitattack: Alleviating intention shift in jailbreak attacks via adaptive bait craft- ing,.

The Scales of Justitia: A Comprehensive Survey on Safety Evaluation of LLMs Baitattack: Alleviating intention shift in jailbreak attacks via adaptive bait craft- ing,

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T10:17:26.615851Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:17:26.615851Z digest=sha256:d3c748ed1685c3484376872488a25df8d7b430b82bc77d0fd5af89bb644433b1

Observation 62e4b8f5-7ef9-451f-8f30-fb596ba32f98 · outbound

This paper cites Feint and Attack: Attention-Based Strategies for Jailbreaking and Protecting LLMs.

The Scales of Justitia: A Comprehensive Survey on Safety Evaluation of LLMs Feint and Attack: Attention-Based Strategies for Jailbreaking and Protecting LLMs

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T10:17:26.618779Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:17:26.618779Z digest=sha256:5d3c5b65f33b7f05014dd062c38f370b347bda8cf65ad075f3a37c26a11010bd

Observation 2b0c7702-ed59-4dc6-85aa-41f05fed7972 · outbound

This paper cites Exploring Vulnerabilities and Protections in Large Language Models: A Survey.

The Scales of Justitia: A Comprehensive Survey on Safety Evaluation of LLMs Exploring Vulnerabilities and Protections in Large Language Models: A Survey

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T10:17:26.621846Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:17:26.621846Z digest=sha256:3278b00b8b46ce9713997e1c5757baa83c0dd8b71e10e9a0af5a46d8c7353fef

Observation d8a0c089-0661-4617-a4b5-db7137f29cea · outbound

This paper cites Evaluating Large Language Models: A Comprehensive Survey.

The Scales of Justitia: A Comprehensive Survey on Safety Evaluation of LLMs Evaluating Large Language Models: A Comprehensive Survey

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T10:17:26.624960Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:17:26.624960Z digest=sha256:374c088f9624fdc44f85b8bd7e7c97b8e91a230181bacd2d01a1d1a4ef706a38

Observation 3c486612-93d0-40dc-85e4-27353eafd775 · outbound

This paper cites Attacks, Defenses and Evaluations for LLM Conversation Safety: A Survey.

The Scales of Justitia: A Comprehensive Survey on Safety Evaluation of LLMs Attacks, Defenses and Evaluations for LLM Conversation Safety: A Survey

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T10:17:26.628023Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:17:26.628023Z digest=sha256:5a73ef1c203818e7dda038e27d1f210b09428125ee4a390b44ee0ad8bbf7dbeb

Observation 77dac2f4-a68d-49be-8a83-9918c71cd095 · outbound

This paper cites Jailbreak Attacks and Defenses Against Large Language Models: A Survey.

The Scales of Justitia: A Comprehensive Survey on Safety Evaluation of LLMs Jailbreak Attacks and Defenses Against Large Language Models: A Survey

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T10:17:26.631200Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:17:26.631200Z digest=sha256:b2fc83dd838e3be4b6c7aeaa641dc8e98dd6e7be17ef72e38b5113f1de3c852f

Observation ff504655-46db-46a2-b087-d917977f14b4 · outbound

This paper cites JailbreakEval: An Integrated Toolkit for Evaluating Jailbreak Attempts Against Large Language Models.

The Scales of Justitia: A Comprehensive Survey on Safety Evaluation of LLMs JailbreakEval: An Integrated Toolkit for Evaluating Jailbreak Attempts Against Large Language Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T10:17:26.634313Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:17:26.634313Z digest=sha256:8224fe14179c3195c807b5580865302a2e98ec706011cb4345d2374db6c79a4a

Observation 2edbefe7-2513-4717-8e2f-d050fbb498da · outbound

This paper cites Safetyprompts: a systematic review of open datasets for evaluating and improving large language model safety,.

The Scales of Justitia: A Comprehensive Survey on Safety Evaluation of LLMs Safetyprompts: a systematic review of open datasets for evaluating and improving large language model safety,

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T10:17:26.637966Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:17:26.637966Z digest=sha256:3ddc2b71723331c12c53132a186da4a524f580173f74ed989e75b315ca8475fa

Observation cabd2fa2-ce3a-46d0-b8eb-b7940e966dde · outbound

This paper cites A Survey of Useful LLM Evaluation.

The Scales of Justitia: A Comprehensive Survey on Safety Evaluation of LLMs A Survey of Useful LLM Evaluation

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T10:17:26.641060Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:17:26.641060Z digest=sha256:0ef1c6acc5abd63bdceb02a264eee59b62b18a6c2c3c5595b91b57436373bea8

Observation efb100cd-e091-4f8a-8418-59cf60dcab4f · outbound

This paper cites ToxiGen: A Large-Scale Machine-Generated Dataset for Adversarial and Implicit Hate Speech Detection.

The Scales of Justitia: A Comprehensive Survey on Safety Evaluation of LLMs ToxiGen: A Large-Scale Machine-Generated Dataset for Adversarial and Implicit Hate Speech Detection

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T10:17:26.644437Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:17:26.644437Z digest=sha256:4089153c90e0663158b4414515bcc8138fe584c4fc667202d214c85ee36fa6ad

Observation 899debdd-8dbc-4652-b50f-8fe1d3245bd9 · outbound

This paper cites ToxicChat: Unveiling Hidden Challenges of Toxicity Detection in Real-World User-AI Conversation.

The Scales of Justitia: A Comprehensive Survey on Safety Evaluation of LLMs ToxicChat: Unveiling Hidden Challenges of Toxicity Detection in Real-World User-AI Conversation

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T10:17:26.647489Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:17:26.647489Z digest=sha256:3b00593451c8669d2cbe2413b31f96c2f8515d2b9813791e43795e6cd3897c8f

Observation df9502d1-5d95-4f67-84f4-d16102d59678 · outbound

This paper cites LifeTox: Unveiling Implicit Toxicity in Life Advice.

The Scales of Justitia: A Comprehensive Survey on Safety Evaluation of LLMs LifeTox: Unveiling Implicit Toxicity in Life Advice

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T10:17:26.651171Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:17:26.651171Z digest=sha256:b33e483211e5d93f06a7f5650aa0bc5526fe58a9dbe95e843064df5d4aaaf973

Observation 86eca9ba-8ef8-4712-bdef-27c0e1b819a1 · outbound

This paper cites PolygloToxicityPrompts: Multilingual Evaluation of Neural Toxic Degeneration in Large Language Models.

The Scales of Justitia: A Comprehensive Survey on Safety Evaluation of LLMs PolygloToxicityPrompts: Multilingual Evaluation of Neural Toxic Degeneration in Large Language Models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T10:17:26.654550Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:17:26.654550Z digest=sha256:6ac646e2df5d2ed3c3a34f70dc6fb3a36d8cd10ae23481ea351dd152ca7a3b19

Observation 73959c27-4331-4ad2-ac15-f280e44a1f53 · outbound

This paper cites Red-Teaming Large Language Models using Chain of Utterances for Safety-Alignment.

The Scales of Justitia: A Comprehensive Survey on Safety Evaluation of LLMs Red-Teaming Large Language Models using Chain of Utterances for Safety-Alignment

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T10:17:26.657917Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:17:26.657917Z digest=sha256:89005ac513d783a8ca70abc6047cb465c24c2e3a7272f6ed450fd83b88771206

Observation f50081d6-ece3-4f50-84f9-57b8b880f014 · outbound

This paper cites FFT: Towards Harmlessness Evaluation and Analysis for LLMs with Factuality, Fairness, Toxicity.

The Scales of Justitia: A Comprehensive Survey on Safety Evaluation of LLMs FFT: Towards Harmlessness Evaluation and Analysis for LLMs with Factuality, Fairness, Toxicity

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T10:17:26.661091Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:17:26.661091Z digest=sha256:5a5a26acccb74ed611c63ecc3dedb5e8ba974fbf0f1aa7218b7940315c5c511a

Observation 0585a788-d1a7-4811-99c0-f81da675aa71 · outbound

This paper cites Red teaming ChatGPT via Jailbreaking: Bias, Robustness, Reliability and Toxicity.

The Scales of Justitia: A Comprehensive Survey on Safety Evaluation of LLMs Red teaming ChatGPT via Jailbreaking: Bias, Robustness, Reliability and Toxicity

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T10:17:26.664466Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:17:26.664466Z digest=sha256:f9d8630ef3060aaaa337c24865d4062fbe234d9ff1da9abbf932801b09a850c0

Observation 2df7b617-89ad-4115-b950-f532132b1e0c · outbound

This paper cites Characterizing and Evaluating the Reliability of LLMs against Jailbreak Attacks.

The Scales of Justitia: A Comprehensive Survey on Safety Evaluation of LLMs Characterizing and Evaluating the Reliability of LLMs against Jailbreak Attacks

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T10:17:26.667702Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:17:26.667702Z digest=sha256:209604be1c382c13ee5143d44e5f8d0ed75e78876924aa03ca5d11946d97374b

Observation 7437ff4a-917b-4a66-8634-2c166b2d0948 · outbound

This paper cites Decodingtrust: A comprehen- sive assessment of trustworthiness in gpt models.

The Scales of Justitia: A Comprehensive Survey on Safety Evaluation of LLMs Decodingtrust: A comprehen- sive assessment of trustworthiness in gpt models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T10:17:26.671111Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:17:26.671111Z digest=sha256:11fb738b2daca68f561180ff4a8b3be64a0a64aafe5325ca7165cf5e6229d2b7

Observation fd341ccf-1802-4f7c-8bd4-19f107d7fd08 · outbound

This paper cites ASSERT: Automated Safety Scenario Red Teaming for Evaluating the Robustness of Large Language Models.

The Scales of Justitia: A Comprehensive Survey on Safety Evaluation of LLMs ASSERT: Automated Safety Scenario Red Teaming for Evaluating the Robustness of Large Language Models

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T10:17:26.674036Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:17:26.674036Z digest=sha256:7a72519b56058093378a956898e527a9acb14c4b07241749b304c4057144a2f6

Observation 246dee32-faa0-4b6a-9816-3e488e486aa4 · outbound

This paper cites Jailbreakbench: An open robustness benchmark for jailbreaking large language models,.

The Scales of Justitia: A Comprehensive Survey on Safety Evaluation of LLMs Jailbreakbench: An open robustness benchmark for jailbreaking large language models,

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T10:17:26.677245Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:17:26.677245Z digest=sha256:2c1d94f7d86c06c2c8b1d59e5069bb34fb889fee43c6aeb89a4e1b1da6a81ec7

Observation bad267d3-7a60-4952-bd1b-c2613b7637b3 · outbound

This paper cites HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust Refusal.

The Scales of Justitia: A Comprehensive Survey on Safety Evaluation of LLMs HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust Refusal

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T10:17:26.680061Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:17:26.680061Z digest=sha256:b6b33730a2d2b1db56e43cb887b81194d887643e27fdc3569187d850a698daac

Observation a65bfd8b-b3c0-4bce-9709-2d22f65ccfec · outbound

This paper cites XSTest: A Test Suite for Identifying Exaggerated Safety Behaviours in Large Language Models.

The Scales of Justitia: A Comprehensive Survey on Safety Evaluation of LLMs XSTest: A Test Suite for Identifying Exaggerated Safety Behaviours in Large Language Models

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T10:17:26.683728Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:17:26.683728Z digest=sha256:4f7145dfbbcbc7fa626ca6fbd6415eade418851e48c633c9c9d8f82cd8913d17

Observation 4bf1d9f4-6e4d-43e9-ae6a-df2c055a56fc · outbound

This paper cites Automatic Pseudo-Harmful Prompt Generation for Evaluating False Refusals in Large Language Models.

The Scales of Justitia: A Comprehensive Survey on Safety Evaluation of LLMs Automatic Pseudo-Harmful Prompt Generation for Evaluating False Refusals in Large Language Models

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T10:17:26.686789Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:17:26.686789Z digest=sha256:a2b58799f4148b6be9399e3cd898fb47d9a70f26735ee2474dcab4dee39cf719

Observation a35ca690-ee7d-4ed1-9390-0c435e34c5e2 · outbound

This paper cites The Art of Defending: A Systematic Evaluation and Analysis of LLM Defense Strategies on Safety and Over-Defensiveness.

The Scales of Justitia: A Comprehensive Survey on Safety Evaluation of LLMs The Art of Defending: A Systematic Evaluation and Analysis of LLM Defense Strategies on Safety and Over-Defensiveness

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T10:17:26.690689Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:17:26.690689Z digest=sha256:33ed3141b76ee26897848062abed3c0fbd629d670e6b6897d752a14814eb0aee

Observation bd275b6f-cf16-4016-84a6-b0e451a0f417 · outbound

This paper cites How (un)ethical are instruction-centric responses of LLMs? Unveiling the vulnerabilities of safety guardrails to harmful queries.

The Scales of Justitia: A Comprehensive Survey on Safety Evaluation of LLMs How (un)ethical are instruction-centric responses of LLMs? Unveiling the vulnerabilities of safety guardrails to harmful queries

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T10:17:26.693976Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:17:26.693976Z digest=sha256:0bcfbc565f58a428005533e4ad88286779f89baec8dbd604d34181f620ebb47d

Observation 12cc8f6e-20cc-46c9-b7c2-170b918e2bf8 · outbound

This paper cites LocalValueBench: A Collaboratively Built and Extensible Benchmark for Evaluating Localized Value Alignment and Ethical Safety in Large Language Models.

The Scales of Justitia: A Comprehensive Survey on Safety Evaluation of LLMs LocalValueBench: A Collaboratively Built and Extensible Benchmark for Evaluating Localized Value Alignment and Ethical Safety in Large Language Models

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T10:17:26.697138Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:17:26.697138Z digest=sha256:1459ce55b4e31232627c08d4ab30b9a3507529bbf97ec7b2ec690a7edf55d7f8

Observation c0e0969a-7990-4d95-86d8-0c75ce9bdd7a · outbound

This paper cites MoralBench: Moral Evaluation of LLMs.

The Scales of Justitia: A Comprehensive Survey on Safety Evaluation of LLMs MoralBench: Moral Evaluation of LLMs

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T10:17:26.700389Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:17:26.700389Z digest=sha256:722529de76276554879742f508ef19a7efbbbd9cd5dfbb69c1d2a0815501a410

Observation 5e5cd74e-3be7-4dbe-abba-1a8aedd5252c · outbound

This paper cites CValues: Measuring the Values of Chinese Large Language Models from Safety to Responsibility.

The Scales of Justitia: A Comprehensive Survey on Safety Evaluation of LLMs CValues: Measuring the Values of Chinese Large Language Models from Safety to Responsibility

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T10:17:26.703758Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:17:26.703758Z digest=sha256:0b4d9b4f3f333980a32b6715552bf2159cfda197a8e91802d3faa7a1717be984

Observation 86fa3827-6a84-4057-addf-83750dbedcec · outbound

This paper cites Safety Assessment of Chinese Large Language Models.

The Scales of Justitia: A Comprehensive Survey on Safety Evaluation of LLMs Safety Assessment of Chinese Large Language Models

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T10:17:26.706992Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:17:26.706992Z digest=sha256:6199dfebc81997b2916f1708d9f19d2b3c63607ce39d60e8e3e878bc9b58489e

Observation 18bd0858-bb54-4480-96c5-5a4b79ac9dd9 · outbound

This paper cites Aligning AI With Shared Human Values.

The Scales of Justitia: A Comprehensive Survey on Safety Evaluation of LLMs Aligning AI With Shared Human Values

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T10:17:26.710261Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:17:26.710261Z digest=sha256:8280f8ac912cfe4cef3845463bc171e22706bf842717630e5f381e9083f71528

Observation 198b5e76-ea6c-46ad-95b0-36c8805787b9 · outbound

This paper cites Are Large Language Models Really Bias-Free? Jailbreak Prompts for Assessing Adversarial Robustness to Bias Elicitation.

The Scales of Justitia: A Comprehensive Survey on Safety Evaluation of LLMs Are Large Language Models Really Bias-Free? Jailbreak Prompts for Assessing Adversarial Robustness to Bias Elicitation

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T10:17:26.713424Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:17:26.713424Z digest=sha256:e11dbf05ae9b621e565aa035d57c45e906cdf9ccbe6f8259cb19bedd8eb09aae

Observation 0ec36478-e17c-4e3b-9364-1a256ec128a5 · outbound

This paper cites Do Large Language Models Rank Fairly? An Empirical Study on the Fairness of LLMs as Rankers.

The Scales of Justitia: A Comprehensive Survey on Safety Evaluation of LLMs Do Large Language Models Rank Fairly? An Empirical Study on the Fairness of LLMs as Rankers

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T10:17:26.716770Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:17:26.716770Z digest=sha256:c5f6c64772c5bb180acb26be68b903081ea0615f1e8414cf5fa07967c5e2f507

Observation 55a85855-b2ce-4594-8d40-04ba76c817d6 · outbound

This paper cites CEB: Compositional Evaluation Benchmark for Fairness in Large Language Models.

The Scales of Justitia: A Comprehensive Survey on Safety Evaluation of LLMs CEB: Compositional Evaluation Benchmark for Fairness in Large Language Models

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T10:17:26.720062Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:17:26.720062Z digest=sha256:536de2391954d61d9a7f9721f60ed1d4ea4cff1f486011400dffe19b65ea24b6

Observation 719252a1-2bff-4c05-aa12-071d48ef29e7 · outbound

This paper cites CBBQ: A Chinese Bias Benchmark Dataset Curated with Human-AI Collaboration for Large Language Models.

The Scales of Justitia: A Comprehensive Survey on Safety Evaluation of LLMs CBBQ: A Chinese Bias Benchmark Dataset Curated with Human-AI Collaboration for Large Language Models

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T10:17:26.723290Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:17:26.723290Z digest=sha256:eac08c60a71e20001a661bac676fa8b649bcab6d1a57b81a6bde0d9374397d29

Observation 4ae11f88-b433-446d-80f2-53ab15d3184e · outbound

This paper cites Llm-driven robots risk enacting discrimination, violence, and unlawful actions,.

The Scales of Justitia: A Comprehensive Survey on Safety Evaluation of LLMs Llm-driven robots risk enacting discrimination, violence, and unlawful actions,

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T10:17:26.726445Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:17:26.726445Z digest=sha256:a0a20a3c2ec53b580fefe2ad2b22bde537e2a7af9184e5f01fa208f30db61876

Observation 0c083f40-2499-4e1c-a67d-2e875657e0cc · outbound

This paper cites Large Language Models are not Fair Evaluators.

The Scales of Justitia: A Comprehensive Survey on Safety Evaluation of LLMs Large Language Models are not Fair Evaluators

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T10:17:26.729392Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:17:26.729392Z digest=sha256:948c2a7b9c2c1de3ae90c1c39e36134688087b6fe8b022ef10eef716e91b168b

Observation 665b996d-3227-477d-8604-74fa96cfbbc2 · outbound

This paper cites How are llms mitigating stereotyping harms? learning from search engine studies,.

The Scales of Justitia: A Comprehensive Survey on Safety Evaluation of LLMs How are llms mitigating stereotyping harms? learning from search engine studies,

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T10:17:26.732664Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:17:26.732664Z digest=sha256:73a6fc49c7b6fd4f0fe68cae444324fac7e8cb69dd0cc75a0b782f680ea79498

Observation 39cbd2d9-dab3-492f-a284-47963571fed9 · outbound

This paper cites AI-LieDar: Examine the Trade-off Between Utility and Truthfulness in LLM Agents.

The Scales of Justitia: A Comprehensive Survey on Safety Evaluation of LLMs AI-LieDar: Examine the Trade-off Between Utility and Truthfulness in LLM Agents

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T10:17:26.735758Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:17:26.735758Z digest=sha256:c99892b50b98105a6a789254fa4d9e734e0c45a40dbdf824aee537448c87b93e

Observation 7cde5061-c429-4817-b56b-05cde472faeb · outbound

This paper cites THaMES: An End-to-End Tool for Hallucination Mitigation and Evaluation in Large Language Models.

The Scales of Justitia: A Comprehensive Survey on Safety Evaluation of LLMs THaMES: An End-to-End Tool for Hallucination Mitigation and Evaluation in Large Language Models

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T10:17:26.738790Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:17:26.738790Z digest=sha256:dad64513f870e947888c3d79a1939355bbfe316b409bbdaf96e78e85130220c1

Observation a9d9be6e-fa23-4842-b1b6-02c01382487d · outbound

This paper cites Reeval: Automatic hallucination evaluation for retrieval-augmented large language models via transferable adversarial attacks,.

The Scales of Justitia: A Comprehensive Survey on Safety Evaluation of LLMs Reeval: Automatic hallucination evaluation for retrieval-augmented large language models via transferable adversarial attacks,

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T10:17:26.742248Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:17:26.742248Z digest=sha256:79f8d9c8e046928ffef5ac385a455926251d73639ab838091c6309670a118cec

Observation ac6ba528-14f1-40f3-8587-189ec7a0cb6d · outbound

This paper cites HaluEval: A Large-Scale Hallucination Evaluation Benchmark for Large Language Models.

The Scales of Justitia: A Comprehensive Survey on Safety Evaluation of LLMs HaluEval: A Large-Scale Hallucination Evaluation Benchmark for Large Language Models

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-07T10:17:26.745240Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:17:26.745240Z digest=sha256:32ae2d7e149247607a37eae7622d86d0811340ea418b09d48021870ee459452c

Observation a939db54-6ef6-4da4-abbb-1433c861881d · outbound

This paper cites Openfactcheck: A unified framework for factuality evaluation of llms,.

The Scales of Justitia: A Comprehensive Survey on Safety Evaluation of LLMs Openfactcheck: A unified framework for factuality evaluation of llms,

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-07T10:17:26.748537Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:17:26.748537Z digest=sha256:d93283eec61aa562476b91d4d4162f0602df4f5d986ec8479b15175f8e215299

Observation b3f68e55-99b5-41d3-abc2-d4854aa90b2b · outbound

This paper cites GraphEval: A Knowledge-Graph Based LLM Hallucination Evaluation Framework.

The Scales of Justitia: A Comprehensive Survey on Safety Evaluation of LLMs GraphEval: A Knowledge-Graph Based LLM Hallucination Evaluation Framework

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-07T10:17:26.751297Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:17:26.751297Z digest=sha256:e6856a51e095e527fd16d656439e8be24c15fe40eb88acb90ba6ea7204b94e1b

Observation fabd4c46-bc56-4d00-8869-3f6d6c3d1eaf · outbound

This paper cites KG-FPQ: Evaluating Factuality Hallucination in LLMs with Knowledge Graph-based False Premise Questions.

The Scales of Justitia: A Comprehensive Survey on Safety Evaluation of LLMs KG-FPQ: Evaluating Factuality Hallucination in LLMs with Knowledge Graph-based False Premise Questions

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-07T10:17:26.754339Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:17:26.754339Z digest=sha256:16f29dff5c27f919db610e09dcf5bec499add0ccee7a0443bb62ef737651eb5b

Observation e02cde99-356d-4ff2-be5e-a779481382e2 · outbound

This paper cites LLM-PBE: Assessing Data Privacy in Large Language Models.

The Scales of Justitia: A Comprehensive Survey on Safety Evaluation of LLMs LLM-PBE: Assessing Data Privacy in Large Language Models

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-07T10:17:26.757454Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:17:26.757454Z digest=sha256:f7f02e879fd03481a033ed66375f571eb9a15739bad2928281b2179ec75f4a03

Observation 01061d32-ffd0-4517-a929-7490bfdff1c3 · outbound

This paper cites MedSafetyBench: Evaluating and Improving the Medical Safety of Large Language Models.

The Scales of Justitia: A Comprehensive Survey on Safety Evaluation of LLMs MedSafetyBench: Evaluating and Improving the Medical Safety of Large Language Models

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-07T10:17:26.760644Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:17:26.760644Z digest=sha256:fc77d85d98355f5629e1805a791713b10196841ad64df170517c35bd26bacd18

Observation 78188daa-cfd8-411e-b556-0bf09ef0c4f5 · outbound

This paper cites Purple Llama CyberSecEval: A Secure Coding Benchmark for Language Models.

The Scales of Justitia: A Comprehensive Survey on Safety Evaluation of LLMs Purple Llama CyberSecEval: A Secure Coding Benchmark for Language Models

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-07T10:17:26.763796Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:17:26.763796Z digest=sha256:bd2d60810ccf317f2e07878a04b815df34fa2a3cad29532abb5b9574bf09f69d

Observation a5c71bf5-c1c4-4f94-841b-1543d665cf28 · outbound

This paper cites Cweval: Outcome- driven evaluation on functionality and security of llm code generation,.

The Scales of Justitia: A Comprehensive Survey on Safety Evaluation of LLMs Cweval: Outcome- driven evaluation on functionality and security of llm code generation,

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-07T10:17:26.766976Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:17:26.766976Z digest=sha256:f4440929495de810755fcbfd7ab039e655e246f16e6f6d40b81a50be55ff8863

Observation 61f9837f-bff1-4280-b78e-f6453c9785b0 · outbound

This paper cites SafeGenBench: A Benchmark Framework for Security Vulnerability Detection in LLM-Generated Code.

The Scales of Justitia: A Comprehensive Survey on Safety Evaluation of LLMs SafeGenBench: A Benchmark Framework for Security Vulnerability Detection in LLM-Generated Code

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-07T10:17:26.769794Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:17:26.769794Z digest=sha256:2bb8a0b16aabfca0b4a7f7c6eff3340287d113567723b6e583cba66a9cb650a3

Observation af3df8c1-40a0-46ea-be95-24be8b49e5bd · outbound

This paper cites SHIELD: Evaluation and Defense Strategies for Copyright Compliance in LLM Text Generation.

The Scales of Justitia: A Comprehensive Survey on Safety Evaluation of LLMs SHIELD: Evaluation and Defense Strategies for Copyright Compliance in LLM Text Generation

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-07T10:17:26.773303Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:17:26.773303Z digest=sha256:63fcc7205dd231aa03811e4b4235e4e74b8d9b2de7bd88039f974150327c2b6f

Observation 0dda50f5-b5d6-496b-9d01-6bbc18b3d89b · outbound

This paper cites TRIDENT: Benchmarking LLM Safety in Finance, Medicine, and Law.

The Scales of Justitia: A Comprehensive Survey on Safety Evaluation of LLMs TRIDENT: Benchmarking LLM Safety in Finance, Medicine, and Law

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-07T10:17:26.776487Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:17:26.776487Z digest=sha256:3c896aed6d0d29df788719dd2fbee50b47fc7c7a8876a2f0c4e3d9133d6b77e9

Observation 523e7bf9-488c-4f03-955d-551ed5e6783d · outbound

This paper cites Rethinking How to Evaluate Language Model Jailbreak.

The Scales of Justitia: A Comprehensive Survey on Safety Evaluation of LLMs Rethinking How to Evaluate Language Model Jailbreak

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-07T10:17:26.779684Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:17:26.779684Z digest=sha256:04f50567c7cd69e0d8d232fc74d58c3569f72c1ec9e434de7d92af388d4876e8

Observation 02c0fdab-1e14-4b37-9335-3bfda1829338 · outbound

This paper cites A StrongREJECT for Empty Jailbreaks.

The Scales of Justitia: A Comprehensive Survey on Safety Evaluation of LLMs A StrongREJECT for Empty Jailbreaks

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-07T10:17:26.782747Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:17:26.782747Z digest=sha256:70d2dade7470852aac684c3518d338e0af1861e18d796fa7375a0205ba366605

Observation 6a1bdefb-1b64-4581-99ed-cd19080fb6c1 · outbound

This paper cites "Not Aligned" is Not "Malicious": Being Careful about Hallucinations of Large Language Models' Jailbreak.

The Scales of Justitia: A Comprehensive Survey on Safety Evaluation of LLMs "Not Aligned" is Not "Malicious": Being Careful about Hallucinations of Large Language Models' Jailbreak

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-07T10:17:26.785868Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:17:26.785868Z digest=sha256:c58a59587e59c42b005e11ad14e654c5fdbcc72464cd2d5599681f47116537d9

Observation 266333ba-7c41-48f7-85d0-e27a04f77d93 · outbound

This paper cites JailbreakRadar: Comprehensive Assessment of Jailbreak Attacks Against LLMs.

The Scales of Justitia: A Comprehensive Survey on Safety Evaluation of LLMs JailbreakRadar: Comprehensive Assessment of Jailbreak Attacks Against LLMs

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-07T10:17:26.788943Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:17:26.788943Z digest=sha256:3ec3e1bf6119ec9be9c033a5a0708c139aa2ece939013e109f3b5e283c65396d

Observation c9e94feb-06c5-456c-9bf8-bdcfb066834b · outbound

This paper cites R-Judge: Benchmarking Safety Risk Awareness for LLM Agents.

The Scales of Justitia: A Comprehensive Survey on Safety Evaluation of LLMs R-Judge: Benchmarking Safety Risk Awareness for LLM Agents

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-07T10:17:26.792352Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:17:26.792352Z digest=sha256:fbf2c7dbf3de4e63c132a1dcef164afffcb23bf0584b2707c64249fba0c9f124

Observation 18774a35-b8c8-4437-ab17-31ec03d9fba4 · outbound

This paper cites Do-not-answer: Evaluating safeguards in llms,.

The Scales of Justitia: A Comprehensive Survey on Safety Evaluation of LLMs Do-not-answer: Evaluating safeguards in llms,

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-07T10:17:26.795573Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:17:26.795573Z digest=sha256:4451f033218fb31f1169eb4e8309ef50e018ca502aa9f9a46035bfcaac7fb024

Observation 57447a39-ddb6-4026-9b68-e17685ee28aa · outbound

This paper cites CoSafe: Evaluating Large Language Model Safety in Multi-Turn Dialogue Coreference.

The Scales of Justitia: A Comprehensive Survey on Safety Evaluation of LLMs CoSafe: Evaluating Large Language Model Safety in Multi-Turn Dialogue Coreference

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-07T10:17:26.798413Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:17:26.798413Z digest=sha256:7038ed6e1acc4b62bf8ee26e5ee60d1745bec22b262c6ddfead6336353398031

Observation 39f87ec4-b572-468a-8ca6-49826f53a9bd · outbound

This paper cites JAILJUDGE: A Comprehensive Jailbreak Judge Benchmark with Multi-Agent Enhanced Explanation Evaluation Framework.

The Scales of Justitia: A Comprehensive Survey on Safety Evaluation of LLMs JAILJUDGE: A Comprehensive Jailbreak Judge Benchmark with Multi-Agent Enhanced Explanation Evaluation Framework

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-07T10:17:26.801585Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:17:26.801585Z digest=sha256:8f423ab6ed4351356c7a4390e67c22ecd9a168e183bb598f59c50b48ad73b7d5

Observation 5d669970-5c06-4d74-98f9-d4fed5a5a342 · outbound

This paper cites SORRY-Bench: Systematically Evaluating Large Language Model Safety Refusal.

The Scales of Justitia: A Comprehensive Survey on Safety Evaluation of LLMs SORRY-Bench: Systematically Evaluating Large Language Model Safety Refusal

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-07T10:17:26.804783Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:17:26.804783Z digest=sha256:a322f1b7e7f01f5635d9f4b87aee3ebdaf698099d49b0293f84b89c62c851da9

Observation ce2e4144-2beb-4ad1-a941-531422c683b6 · outbound

This paper cites AEGIS: Online Adaptive AI Content Safety Moderation with Ensemble of LLM Experts.

The Scales of Justitia: A Comprehensive Survey on Safety Evaluation of LLMs AEGIS: Online Adaptive AI Content Safety Moderation with Ensemble of LLM Experts

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-07T10:17:26.807876Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:17:26.807876Z digest=sha256:6138d5e946c5d93e2bd69caa5c74e5bf74a7b56e88e606e7fe7c15b9c18779e0

Observation d7eb6bf1-2690-4630-8596-899c53113dab · outbound

This paper cites SALAD-Bench: A Hierarchical and Comprehensive Safety Benchmark for Large Language Models.

The Scales of Justitia: A Comprehensive Survey on Safety Evaluation of LLMs SALAD-Bench: A Hierarchical and Comprehensive Safety Benchmark for Large Language Models

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-07T10:17:26.810690Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:17:26.810690Z digest=sha256:df3bdf900fe1ed29b943fce49181fe74e8b07ef3ddfd4bb3b95cca8f2b58f045

Observation 572e6694-eddc-49d5-851b-3534c1deaa90 · outbound

This paper cites ALERT: A Comprehensive Benchmark for Assessing Large Language Models' Safety through Red Teaming.

The Scales of Justitia: A Comprehensive Survey on Safety Evaluation of LLMs ALERT: A Comprehensive Benchmark for Assessing Large Language Models' Safety through Red Teaming

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-07T10:17:26.813755Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:17:26.813755Z digest=sha256:c8b09e5bd87549c681bd1841568eaa44905682a6cf2d029489d0abaff76bd04b

Observation cf1de549-be79-4d74-a134-3f4e255ace61 · outbound

This paper cites WildGuard: Open One-Stop Moderation Tools for Safety Risks, Jailbreaks, and Refusals of LLMs.

The Scales of Justitia: A Comprehensive Survey on Safety Evaluation of LLMs WildGuard: Open One-Stop Moderation Tools for Safety Risks, Jailbreaks, and Refusals of LLMs

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-07T10:17:26.817179Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:17:26.817179Z digest=sha256:eca6da9d7037fb07a01bf78d33642659b3a410f7f8e189d6d302dc0f52d9bee6

Observation 8ba49cd3-3476-4039-9bd9-4f94ab3c2c62 · outbound

This paper cites "Do Anything Now": Characterizing and Evaluating In-The-Wild Jailbreak Prompts on Large Language Models.

The Scales of Justitia: A Comprehensive Survey on Safety Evaluation of LLMs "Do Anything Now": Characterizing and Evaluating In-The-Wild Jailbreak Prompts on Large Language Models

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-07T10:17:26.820201Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:17:26.820201Z digest=sha256:141fe21bb50ce93b0f4ab0d2fe76388ba8b9b52fec69609989b8ac47628ceaba

Observation add6c427-c3f2-48e9-81f7-29d979303741 · outbound

This paper cites CHiSafetyBench: A Chinese Hierarchical Safety Benchmark for Large Language Models.

The Scales of Justitia: A Comprehensive Survey on Safety Evaluation of LLMs CHiSafetyBench: A Chinese Hierarchical Safety Benchmark for Large Language Models

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-07T10:17:26.823250Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:17:26.823250Z digest=sha256:9924ee8f673c52802785d7a1bade6b5f0374f943d000cc8647b4465adc188465

Observation 5b538f84-aa5b-4c16-8f9c-b01e0a256f96 · outbound

This paper cites A Chinese Dataset for Evaluating the Safeguards in Large Language Models.

The Scales of Justitia: A Comprehensive Survey on Safety Evaluation of LLMs A Chinese Dataset for Evaluating the Safeguards in Large Language Models

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-07T10:17:26.826456Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:17:26.826456Z digest=sha256:e196185efa309c9e7bbf20b9871a755ceb311486cab27627a7d5795a6d935dd7

Observation 5e22d829-7044-4107-85cd-2be2c5548390 · outbound

This paper cites JailBench: A Comprehensive Chinese Security Assessment Benchmark for Large Language Models.

The Scales of Justitia: A Comprehensive Survey on Safety Evaluation of LLMs JailBench: A Comprehensive Chinese Security Assessment Benchmark for Large Language Models

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-07T10:17:26.829492Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:17:26.829492Z digest=sha256:783ed05aa9fd6a027bf4ccfe93ca0a381ec3ab500378d8a72d54d5fd97da97aa

Observation ca731799-ea27-4e7a-9c7b-1f1e839cc98a · outbound

This paper cites CRiskEval: A Chinese Multi-Level Risk Evaluation Benchmark Dataset for Large Language Models.

The Scales of Justitia: A Comprehensive Survey on Safety Evaluation of LLMs CRiskEval: A Chinese Multi-Level Risk Evaluation Benchmark Dataset for Large Language Models

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-07T10:17:26.832528Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:17:26.832528Z digest=sha256:b9424d3546c5b633d72574f623459d89d98a22e047651771f542a7ab24af8f12

Observation 5ff719d5-d45f-4af8-8b86-829aab8d9f95 · outbound

This paper cites ChineseSafe: A Chinese Benchmark for Evaluating Safety in Large Language Models.

The Scales of Justitia: A Comprehensive Survey on Safety Evaluation of LLMs ChineseSafe: A Chinese Benchmark for Evaluating Safety in Large Language Models

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-07T10:17:26.835560Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:17:26.835560Z digest=sha256:28883b8fe9d1a7a18a4aa6f21bcfc1d6c8d0cc7a8c64b4a9eb77114b556a4086

Observation 4d02dc49-52e6-4edc-bd31-c0d406093cc2 · outbound

This paper cites Latent Jailbreak: A Benchmark for Evaluating Text Safety and Output Robustness of Large Language Models.

The Scales of Justitia: A Comprehensive Survey on Safety Evaluation of LLMs Latent Jailbreak: A Benchmark for Evaluating Text Safety and Output Robustness of Large Language Models

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-07T10:17:26.838444Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:17:26.838444Z digest=sha256:658051a70910abeecf1ffb1e69aa82425086b5222798ef33c0464072c455f663

Observation b06fc21c-6af4-4cd9-8846-993cfd87ade3 · outbound

This paper cites Safetybench: Evaluating the safety of large language models,.

The Scales of Justitia: A Comprehensive Survey on Safety Evaluation of LLMs Safetybench: Evaluating the safety of large language models,

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-07T10:17:26.841712Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:17:26.841712Z digest=sha256:70a4ae7ed36a6c312319f9d93e4cb37a2bbe73459b6fcc564875fad1d8151123

Observation d4ac6a66-df10-4627-9818-94933ba03a0e · outbound

This paper cites S-Eval: Towards Automated and Comprehensive Safety Evaluation for Large Language Models.

The Scales of Justitia: A Comprehensive Survey on Safety Evaluation of LLMs S-Eval: Towards Automated and Comprehensive Safety Evaluation for Large Language Models

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-07T10:17:26.844499Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:17:26.844499Z digest=sha256:d4fe1828316a0124b004d7b9fe81c7b60f11b85f365622f83e602db113b016c6

Observation ba7ea828-025e-419d-8261-7f473360d64b · outbound

This paper cites AraTrust: An Evaluation of Trustworthiness for LLMs in Arabic.

The Scales of Justitia: A Comprehensive Survey on Safety Evaluation of LLMs AraTrust: An Evaluation of Trustworthiness for LLMs in Arabic

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-07T10:17:26.847859Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:17:26.847859Z digest=sha256:810efdf2201a97327f1194f1ff13e608023050042f471a711bdea6a6f8e3489b

Observation 5cf06d1a-4497-4622-81af-3c946d84e8ef · outbound

This paper cites AnswerCarefully: A Dataset for Improving the Safety of Japanese LLM Output.

The Scales of Justitia: A Comprehensive Survey on Safety Evaluation of LLMs AnswerCarefully: A Dataset for Improving the Safety of Japanese LLM Output

Reference 85

Resolution
unresolved
no resolver link, observed 2026-08-07T10:17:26.850823Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:17:26.850823Z digest=sha256:1cd86ec3fc3934da3b6febf20e6de06b341be7fc959a3a776f5b17241b83acc0

Observation a560036a-b759-40a8-94fa-e2463d811690 · outbound

This paper cites All languages matter: On the multilingual safety of llms,.

The Scales of Justitia: A Comprehensive Survey on Safety Evaluation of LLMs All languages matter: On the multilingual safety of llms,

Reference 86

Resolution
unresolved
no resolver link, observed 2026-08-07T10:17:26.853966Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:17:26.853966Z digest=sha256:3a5d52172ea1b199949cdb689c71bac95c6e71e9f273b460bc6f737bc5a6fdf9

Observation 0a192469-b02d-4b1a-8cf1-e5ce6f13696e · outbound

This paper cites Annotation alignment: Comparing LLM and human annotations of conversational safety.

The Scales of Justitia: A Comprehensive Survey on Safety Evaluation of LLMs Annotation alignment: Comparing LLM and human annotations of conversational safety

Reference 87

Resolution
unresolved
no resolver link, observed 2026-08-07T10:17:26.856833Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:17:26.856833Z digest=sha256:306584de31e63d76a701494b77858788e1a2533b2ea207804228c98aea9854f1

Observation d69db30b-4bb4-4c01-aa10-698b3e24da60 · outbound

This paper cites Making them ask and answer: Jailbreaking large language models in few queries via disguise and reconstruction,.

The Scales of Justitia: A Comprehensive Survey on Safety Evaluation of LLMs Making them ask and answer: Jailbreaking large language models in few queries via disguise and reconstruction,

Reference 88

Resolution
unresolved
no resolver link, observed 2026-08-07T10:17:26.859660Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:17:26.859660Z digest=sha256:2f8407521cf1fd5fddd5743b1958fe913c320d1e42c271dddcd492c1797e36c7

Observation 20c6fe58-3282-48ed-ad79-6b3d4083e531 · outbound

This paper cites AutoDefense: Multi-Agent LLM Defense against Jailbreak Attacks.

The Scales of Justitia: A Comprehensive Survey on Safety Evaluation of LLMs AutoDefense: Multi-Agent LLM Defense against Jailbreak Attacks

Reference 89

Resolution
unresolved
no resolver link, observed 2026-08-07T10:17:26.862405Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:17:26.862405Z digest=sha256:924893d2682584a89f5fcf9d90be22a5c505aed0aa150a5acb2c77d5fe2f5dbb

Observation 59200a37-2356-4583-acc2-ca972c923fea · outbound

This paper cites OpenAI Moderation Endpoint,.

The Scales of Justitia: A Comprehensive Survey on Safety Evaluation of LLMs OpenAI Moderation Endpoint,

Reference 90

Resolution
unresolved
no resolver link, observed 2026-08-07T10:17:26.865748Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:17:26.865748Z digest=sha256:0024b39484050dfea0ea931a79daed4ae45ba31262f17658611bc57ccaee7498

Observation 7a5329e4-ad5b-4d08-aa82-9428ead862d0 · outbound

This paper cites Google Perspective API,.

The Scales of Justitia: A Comprehensive Survey on Safety Evaluation of LLMs Google Perspective API,

Reference 91

Resolution
unresolved
no resolver link, observed 2026-08-07T10:17:26.868477Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:17:26.868477Z digest=sha256:4c67987ecafb04c3d188103da69fca570bedff1fff9bfb9098866d444868aa43

Observation 75da5195-c3f9-48c7-8bf0-0314bf98f532 · outbound

This paper cites Microsoft Azure AI Content Safety API,.

The Scales of Justitia: A Comprehensive Survey on Safety Evaluation of LLMs Microsoft Azure AI Content Safety API,

Reference 92

Resolution
unresolved
no resolver link, observed 2026-08-07T10:17:26.871153Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:17:26.871153Z digest=sha256:fc1baa178d10f1f4cb355e60f9b7cf7ffbb2e75555bae7ee051285f731bc570a

Observation ac0c4c18-e839-48df-8d70-87d8e6188ef2 · outbound

This paper cites GPTFUZZER: Red Teaming Large Language Models with Auto-Generated Jailbreak Prompts.

The Scales of Justitia: A Comprehensive Survey on Safety Evaluation of LLMs GPTFUZZER: Red Teaming Large Language Models with Auto-Generated Jailbreak Prompts

Reference 93

Resolution
unresolved
no resolver link, observed 2026-08-07T10:17:26.873729Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:17:26.873729Z digest=sha256:df78b9b36ef814ea1a1ecbdc2c7c877d6aee95646b4f8f56063869a2fc6b7d3e

Observation e18b27fe-1e54-491a-a7c3-874bbcf659ff · outbound

This paper cites Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!.

The Scales of Justitia: A Comprehensive Survey on Safety Evaluation of LLMs Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!

Reference 94

Resolution
unresolved
no resolver link, observed 2026-08-07T10:17:26.876762Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:17:26.876762Z digest=sha256:06b544f92ffe95d33e27874cd05afe9dd603640d2089285b5162846414a605b1

Observation 6627fce0-2607-4a62-b456-df324fc6b2e2 · outbound

This paper cites Jailbreaking Black Box Large Language Models in Twenty Queries.

The Scales of Justitia: A Comprehensive Survey on Safety Evaluation of LLMs Jailbreaking Black Box Large Language Models in Twenty Queries

Reference 95

Resolution
unresolved
no resolver link, observed 2026-08-07T10:17:26.879793Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:17:26.879793Z digest=sha256:00a713d806accab72b1d5b542db218b9316fe2e208f4a28d4d69b62f1427db81

Observation 3c351043-6e22-4ba2-8063-e415567b38df · outbound

This paper cites Scalable and Transferable Black-Box Jailbreaks for Language Models via Persona Modulation.

The Scales of Justitia: A Comprehensive Survey on Safety Evaluation of LLMs Scalable and Transferable Black-Box Jailbreaks for Language Models via Persona Modulation

Reference 96

Resolution
unresolved
no resolver link, observed 2026-08-07T10:17:26.883048Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:17:26.883048Z digest=sha256:10df89aefd23ba7ce38336499c46c2b6d9351adf5cf1f43652a74ee4b2705bf2

Observation 424edd04-e787-49a4-8572-4b0d520bdb7d · outbound

This paper cites Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations.

The Scales of Justitia: A Comprehensive Survey on Safety Evaluation of LLMs Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations

Reference 97

Resolution
unresolved
no resolver link, observed 2026-08-07T10:17:26.886326Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:17:26.886326Z digest=sha256:7b222b043f0bbc257afc04207b60d14f39330b8fda9999614b058989ee986498

Observation fb742edd-25b6-40c3-8c08-6af92546d598 · outbound

This paper cites Meta llama guard 2,.

The Scales of Justitia: A Comprehensive Survey on Safety Evaluation of LLMs Meta llama guard 2,

Reference 98

Resolution
unresolved
no resolver link, observed 2026-08-07T10:17:26.890022Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:17:26.890022Z digest=sha256:c0a728c3118ae7cce21a13a6ac0f51b24804353aca4a16811ce83553c284eaa4

Observation cf752b6e-8bba-46ef-8dac-e08e997d178e · outbound

This paper cites The llama 3 family of models,.

The Scales of Justitia: A Comprehensive Survey on Safety Evaluation of LLMs The llama 3 family of models,

Reference 99

Resolution
unresolved
no resolver link, observed 2026-08-07T10:17:26.893183Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:17:26.893183Z digest=sha256:7829c8886c271e99cd2abdb7c23c47a6a99970b631236c15658c4de4944d5a3c

Observation d80a1f43-2742-44ff-a2a1-31cf88ef4dc3 · outbound

This paper cites The Llama 3 Herd of Models.

The Scales of Justitia: A Comprehensive Survey on Safety Evaluation of LLMs The Llama 3 Herd of Models

Reference 100

Resolution
unresolved
no resolver link, observed 2026-08-07T10:17:26.896104Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:17:26.896104Z digest=sha256:bcf7aee6ede886be48bfafa52ca8d234cea7460d885ee257e974a27b559ec59d

Pith citing papers

Observation 02d67783-e274-46fb-8f3c-c6728334936a · inbound

PRISON: Unmasking the Criminal Potential of Large Language Models cites this paper.

PRISON: Unmasking the Criminal Potential of Large Language Models The Scales of Justitia: A Comprehensive Survey on Safety Evaluation of LLMs

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-06T23:49:46.938526Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:49:46.938526Z digest=sha256:ab8e76056eff346dba00e55c3b5b395002ec3d44888fa7f8da6e4c73be6395ca

Observation 36cd4c5f-04c3-46b8-9ee2-b89c5edef328 · inbound

LLMEval-Fair: A Large-Scale Longitudinal Study on Robust and Fair Evaluation of Large Language Models cites this paper.

LLMEval-Fair: A Large-Scale Longitudinal Study on Robust and Fair Evaluation of Large Language Models The Scales of Justitia: A Comprehensive Survey on Safety Evaluation of LLMs

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-07-01T02:17:15.985374Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-19T00:39:02.015913Z digest=sha256:f7c2ecbb7e59c9675c824fe6555ee61425afd894fddce6509a4e7631772351ba

Observation be0be778-36e8-4acb-b494-9b96d9cca347 · inbound

Safe for Whom? Rethinking How We Evaluate the Safety of LLMs for Real Users cites this paper.

Safe for Whom? Rethinking How We Evaluate the Safety of LLMs for Real Users The Scales of Justitia: A Comprehensive Survey on Safety Evaluation of LLMs

Reference 11

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T02:17:15.985374Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-16T23:22:42.431997Z digest=sha256:0ae54c46916cb47785998ee85e81ae56fe94ca40ce86abe9acb249c31693d785

Observation b5b39288-6849-4912-bf6c-b08044cad947 · inbound

When Safety Becomes a Vulnerability: Exploiting LLM Alignment Homogeneity for Transferable Blocking in RAG cites this paper.

When Safety Becomes a Vulnerability: Exploiting LLM Alignment Homogeneity for Transferable Blocking in RAG The Scales of Justitia: A Comprehensive Survey on Safety Evaluation of LLMs

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-03T01:25:15.252924Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T01:25:15.252924Z digest=sha256:e998cf16f9ddf44760e3332d9980e89ec736c04bc8f97849ad9762bd31057de8

Observation 3ec71ee7-047f-488c-bf7a-63988fe1b47c · inbound

Steering at the Source: Style Modulation Heads for Robust Persona Control cites this paper.

Steering at the Source: Style Modulation Heads for Robust Persona Control The Scales of Justitia: A Comprehensive Survey on Safety Evaluation of LLMs

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-02T21:21:25.716257Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:21:25.716257Z digest=sha256:2891e5794560bfd11a44d94ff7a54ad8becfa1f20db9fd9526386ce4dc787cbc

Observation 0a90883e-cb2f-478e-b566-4c22d751e037 · inbound

A Systematic Evaluation of Black-Box Uncertainty Estimation Methods for Large Language Models cites this paper.

A Systematic Evaluation of Black-Box Uncertainty Estimation Methods for Large Language Models The Scales of Justitia: A Comprehensive Survey on Safety Evaluation of LLMs

Reference 11

Resolution
metadata mismatch
local_arxiv, observed 2026-07-04T03:39:30.629404Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-06-26T17:46:00.822375Z digest=sha256:dfbd631e35bf977c730c5b9ac2cbfc8f3453fb194e5b55c44f613fae88a890a3

Observation 23fe07c8-e95d-47bc-adea-783bcbb938f0 · inbound

Efficient Safety Benchmarking via Item Response Theory cites this paper.

Efficient Safety Benchmarking via Item Response Theory The Scales of Justitia: A Comprehensive Survey on Safety Evaluation of LLMs

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-07-01T15:55:48.959767Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-07-01T15:51:00.484343Z digest=sha256:85635f4f7a3b93078ac78ebf1459336efceb03080d6ee10a6845d0ab59a587b7

Observation 0c4dbdd8-cfb3-482e-9fd6-5dcf00b4372c · inbound

Testing Retrieval-Augmented Generation Systems with Chunk Coverage cites this paper.

Testing Retrieval-Augmented Generation Systems with Chunk Coverage The Scales of Justitia: A Comprehensive Survey on Safety Evaluation of LLMs

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-01T15:53:59.443085Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T15:53:59.443085Z digest=sha256:afdfe121c2746415dca3c5107301c2ede9f201c69e42ef757f6a3fc22b5f0518

Observation d46ec9ac-96fe-416b-9832-b160189c2997 · inbound

Large Language Models and Social Media Information Integrity: Opportunities, Challenges, and Research Directions cites this paper.

Large Language Models and Social Media Information Integrity: Opportunities, Challenges, and Research Directions The Scales of Justitia: A Comprehensive Survey on Safety Evaluation of LLMs

Reference 112

Resolution
unresolved
no resolver link, observed 2026-08-08T18:47:38.262660Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T18:47:38.262660Z digest=sha256:3f64cc6df988d7ee79ffa1495d8221dc531f63eb14272626f80340c79e4c0a29