Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T10:17:26.896104Z
Paper Citation Record · LEDGER
As of 11 August 2026, this Paper Citation Record lists 100 of 215 outbound references and 9 inbound Pith citation observations for arXiv:2506.11094.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T10:17:26.896104Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-11T06:34:44.6726+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-08T18:47:38.262660Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-07-04T03:39:30.627528Z
100 of 215 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 0d671c9f-27dc-4248-92ca-ef486691a2b3 · outbound
The Scales of Justitia: A Comprehensive Survey on Safety Evaluation of LLMs A survey on evaluation of large language models,
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 79aab667-973b-4619-8aae-7c2fb795b132 · outbound
The Scales of Justitia: A Comprehensive Survey on Safety Evaluation of LLMs A systematic review on human and computer interaction,
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7dee63a2-a51e-4896-84cb-287a56989c97 · outbound
The Scales of Justitia: A Comprehensive Survey on Safety Evaluation of LLMs From Handcrafted Features to LLMs: A Brief Survey for Machine Translation Quality Estimation
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 794a941e-1db0-4fa3-8c6f-cdc41b73ba3e · outbound
The Scales of Justitia: A Comprehensive Survey on Safety Evaluation of LLMs A Survey on Evaluating Large Language Models in Code Generation Tasks
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8f2b867d-b91f-44e1-bce2-84835ad69aa5 · outbound
The Scales of Justitia: A Comprehensive Survey on Safety Evaluation of LLMs Beyond Self-Talk: A Communication-Centric Survey of LLM-Based Multi-Agent Systems
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e2cf32aa-d78e-41dc-a8cc-284c5133b36b · outbound
The Scales of Justitia: A Comprehensive Survey on Safety Evaluation of LLMs Unleash LLMs Potential for Recommendation by Coordinating Twin-Tower Dynamic Semantic Token Generator
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3608d722-55aa-45b5-8579-e4e6c896ec9e · outbound
The Scales of Justitia: A Comprehensive Survey on Safety Evaluation of LLMs Beyond Accuracy: Evaluating the Reasoning Behavior of Large Language Models -- A Survey
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3c8b545b-a87d-4e72-8e17-c060b15cf272 · outbound
The Scales of Justitia: A Comprehensive Survey on Safety Evaluation of LLMs Universal and Transferable Adversarial Attacks on Aligned Language Models
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 07958821-5743-4529-b082-170c73e16931 · outbound
The Scales of Justitia: A Comprehensive Survey on Safety Evaluation of LLMs Is chatgpt fair for recommendation? evaluating fairness in large language model recommendation,
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9fc1dd70-6916-4333-b1ca-16117d61bcc0 · outbound
The Scales of Justitia: A Comprehensive Survey on Safety Evaluation of LLMs FineFake: A Knowledge-Enriched Dataset for Fine-Grained Multi-Domain Fake News Detection
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4a85706f-d7d3-4a80-9b79-02a63d2b3316 · outbound
The Scales of Justitia: A Comprehensive Survey on Safety Evaluation of LLMs Baitattack: Alleviating intention shift in jailbreak attacks via adaptive bait craft- ing,
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 62e4b8f5-7ef9-451f-8f30-fb596ba32f98 · outbound
The Scales of Justitia: A Comprehensive Survey on Safety Evaluation of LLMs Feint and Attack: Attention-Based Strategies for Jailbreaking and Protecting LLMs
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2b0c7702-ed59-4dc6-85aa-41f05fed7972 · outbound
The Scales of Justitia: A Comprehensive Survey on Safety Evaluation of LLMs Exploring Vulnerabilities and Protections in Large Language Models: A Survey
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d8a0c089-0661-4617-a4b5-db7137f29cea · outbound
The Scales of Justitia: A Comprehensive Survey on Safety Evaluation of LLMs Evaluating Large Language Models: A Comprehensive Survey
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3c486612-93d0-40dc-85e4-27353eafd775 · outbound
The Scales of Justitia: A Comprehensive Survey on Safety Evaluation of LLMs Attacks, Defenses and Evaluations for LLM Conversation Safety: A Survey
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 77dac2f4-a68d-49be-8a83-9918c71cd095 · outbound
The Scales of Justitia: A Comprehensive Survey on Safety Evaluation of LLMs Jailbreak Attacks and Defenses Against Large Language Models: A Survey
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ff504655-46db-46a2-b087-d917977f14b4 · outbound
The Scales of Justitia: A Comprehensive Survey on Safety Evaluation of LLMs JailbreakEval: An Integrated Toolkit for Evaluating Jailbreak Attempts Against Large Language Models
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2edbefe7-2513-4717-8e2f-d050fbb498da · outbound
The Scales of Justitia: A Comprehensive Survey on Safety Evaluation of LLMs Safetyprompts: a systematic review of open datasets for evaluating and improving large language model safety,
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cabd2fa2-ce3a-46d0-b8eb-b7940e966dde · outbound
The Scales of Justitia: A Comprehensive Survey on Safety Evaluation of LLMs A Survey of Useful LLM Evaluation
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation efb100cd-e091-4f8a-8418-59cf60dcab4f · outbound
The Scales of Justitia: A Comprehensive Survey on Safety Evaluation of LLMs ToxiGen: A Large-Scale Machine-Generated Dataset for Adversarial and Implicit Hate Speech Detection
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 899debdd-8dbc-4652-b50f-8fe1d3245bd9 · outbound
The Scales of Justitia: A Comprehensive Survey on Safety Evaluation of LLMs ToxicChat: Unveiling Hidden Challenges of Toxicity Detection in Real-World User-AI Conversation
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation df9502d1-5d95-4f67-84f4-d16102d59678 · outbound
The Scales of Justitia: A Comprehensive Survey on Safety Evaluation of LLMs LifeTox: Unveiling Implicit Toxicity in Life Advice
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 86eca9ba-8ef8-4712-bdef-27c0e1b819a1 · outbound
The Scales of Justitia: A Comprehensive Survey on Safety Evaluation of LLMs PolygloToxicityPrompts: Multilingual Evaluation of Neural Toxic Degeneration in Large Language Models
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 73959c27-4331-4ad2-ac15-f280e44a1f53 · outbound
The Scales of Justitia: A Comprehensive Survey on Safety Evaluation of LLMs Red-Teaming Large Language Models using Chain of Utterances for Safety-Alignment
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f50081d6-ece3-4f50-84f9-57b8b880f014 · outbound
The Scales of Justitia: A Comprehensive Survey on Safety Evaluation of LLMs FFT: Towards Harmlessness Evaluation and Analysis for LLMs with Factuality, Fairness, Toxicity
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0585a788-d1a7-4811-99c0-f81da675aa71 · outbound
The Scales of Justitia: A Comprehensive Survey on Safety Evaluation of LLMs Red teaming ChatGPT via Jailbreaking: Bias, Robustness, Reliability and Toxicity
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2df7b617-89ad-4115-b950-f532132b1e0c · outbound
The Scales of Justitia: A Comprehensive Survey on Safety Evaluation of LLMs Characterizing and Evaluating the Reliability of LLMs against Jailbreak Attacks
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7437ff4a-917b-4a66-8634-2c166b2d0948 · outbound
The Scales of Justitia: A Comprehensive Survey on Safety Evaluation of LLMs Decodingtrust: A comprehen- sive assessment of trustworthiness in gpt models
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fd341ccf-1802-4f7c-8bd4-19f107d7fd08 · outbound
The Scales of Justitia: A Comprehensive Survey on Safety Evaluation of LLMs ASSERT: Automated Safety Scenario Red Teaming for Evaluating the Robustness of Large Language Models
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 246dee32-faa0-4b6a-9816-3e488e486aa4 · outbound
The Scales of Justitia: A Comprehensive Survey on Safety Evaluation of LLMs Jailbreakbench: An open robustness benchmark for jailbreaking large language models,
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bad267d3-7a60-4952-bd1b-c2613b7637b3 · outbound
The Scales of Justitia: A Comprehensive Survey on Safety Evaluation of LLMs HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust Refusal
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a65bfd8b-b3c0-4bce-9709-2d22f65ccfec · outbound
The Scales of Justitia: A Comprehensive Survey on Safety Evaluation of LLMs XSTest: A Test Suite for Identifying Exaggerated Safety Behaviours in Large Language Models
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4bf1d9f4-6e4d-43e9-ae6a-df2c055a56fc · outbound
The Scales of Justitia: A Comprehensive Survey on Safety Evaluation of LLMs Automatic Pseudo-Harmful Prompt Generation for Evaluating False Refusals in Large Language Models
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a35ca690-ee7d-4ed1-9390-0c435e34c5e2 · outbound
The Scales of Justitia: A Comprehensive Survey on Safety Evaluation of LLMs The Art of Defending: A Systematic Evaluation and Analysis of LLM Defense Strategies on Safety and Over-Defensiveness
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bd275b6f-cf16-4016-84a6-b0e451a0f417 · outbound
The Scales of Justitia: A Comprehensive Survey on Safety Evaluation of LLMs How (un)ethical are instruction-centric responses of LLMs? Unveiling the vulnerabilities of safety guardrails to harmful queries
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 12cc8f6e-20cc-46c9-b7c2-170b918e2bf8 · outbound
The Scales of Justitia: A Comprehensive Survey on Safety Evaluation of LLMs LocalValueBench: A Collaboratively Built and Extensible Benchmark for Evaluating Localized Value Alignment and Ethical Safety in Large Language Models
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c0e0969a-7990-4d95-86d8-0c75ce9bdd7a · outbound
The Scales of Justitia: A Comprehensive Survey on Safety Evaluation of LLMs MoralBench: Moral Evaluation of LLMs
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5e5cd74e-3be7-4dbe-abba-1a8aedd5252c · outbound
The Scales of Justitia: A Comprehensive Survey on Safety Evaluation of LLMs CValues: Measuring the Values of Chinese Large Language Models from Safety to Responsibility
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 86fa3827-6a84-4057-addf-83750dbedcec · outbound
The Scales of Justitia: A Comprehensive Survey on Safety Evaluation of LLMs Safety Assessment of Chinese Large Language Models
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 18bd0858-bb54-4480-96c5-5a4b79ac9dd9 · outbound
The Scales of Justitia: A Comprehensive Survey on Safety Evaluation of LLMs Aligning AI With Shared Human Values
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 198b5e76-ea6c-46ad-95b0-36c8805787b9 · outbound
The Scales of Justitia: A Comprehensive Survey on Safety Evaluation of LLMs Are Large Language Models Really Bias-Free? Jailbreak Prompts for Assessing Adversarial Robustness to Bias Elicitation
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0ec36478-e17c-4e3b-9364-1a256ec128a5 · outbound
The Scales of Justitia: A Comprehensive Survey on Safety Evaluation of LLMs Do Large Language Models Rank Fairly? An Empirical Study on the Fairness of LLMs as Rankers
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 55a85855-b2ce-4594-8d40-04ba76c817d6 · outbound
The Scales of Justitia: A Comprehensive Survey on Safety Evaluation of LLMs CEB: Compositional Evaluation Benchmark for Fairness in Large Language Models
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 719252a1-2bff-4c05-aa12-071d48ef29e7 · outbound
The Scales of Justitia: A Comprehensive Survey on Safety Evaluation of LLMs CBBQ: A Chinese Bias Benchmark Dataset Curated with Human-AI Collaboration for Large Language Models
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4ae11f88-b433-446d-80f2-53ab15d3184e · outbound
The Scales of Justitia: A Comprehensive Survey on Safety Evaluation of LLMs Llm-driven robots risk enacting discrimination, violence, and unlawful actions,
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0c083f40-2499-4e1c-a67d-2e875657e0cc · outbound
The Scales of Justitia: A Comprehensive Survey on Safety Evaluation of LLMs Large Language Models are not Fair Evaluators
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 665b996d-3227-477d-8604-74fa96cfbbc2 · outbound
The Scales of Justitia: A Comprehensive Survey on Safety Evaluation of LLMs How are llms mitigating stereotyping harms? learning from search engine studies,
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 39cbd2d9-dab3-492f-a284-47963571fed9 · outbound
The Scales of Justitia: A Comprehensive Survey on Safety Evaluation of LLMs AI-LieDar: Examine the Trade-off Between Utility and Truthfulness in LLM Agents
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7cde5061-c429-4817-b56b-05cde472faeb · outbound
The Scales of Justitia: A Comprehensive Survey on Safety Evaluation of LLMs THaMES: An End-to-End Tool for Hallucination Mitigation and Evaluation in Large Language Models
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a9d9be6e-fa23-4842-b1b6-02c01382487d · outbound
The Scales of Justitia: A Comprehensive Survey on Safety Evaluation of LLMs Reeval: Automatic hallucination evaluation for retrieval-augmented large language models via transferable adversarial attacks,
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ac6ba528-14f1-40f3-8587-189ec7a0cb6d · outbound
The Scales of Justitia: A Comprehensive Survey on Safety Evaluation of LLMs HaluEval: A Large-Scale Hallucination Evaluation Benchmark for Large Language Models
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a939db54-6ef6-4da4-abbb-1433c861881d · outbound
The Scales of Justitia: A Comprehensive Survey on Safety Evaluation of LLMs Openfactcheck: A unified framework for factuality evaluation of llms,
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b3f68e55-99b5-41d3-abc2-d4854aa90b2b · outbound
The Scales of Justitia: A Comprehensive Survey on Safety Evaluation of LLMs GraphEval: A Knowledge-Graph Based LLM Hallucination Evaluation Framework
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fabd4c46-bc56-4d00-8869-3f6d6c3d1eaf · outbound
The Scales of Justitia: A Comprehensive Survey on Safety Evaluation of LLMs KG-FPQ: Evaluating Factuality Hallucination in LLMs with Knowledge Graph-based False Premise Questions
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e02cde99-356d-4ff2-be5e-a779481382e2 · outbound
The Scales of Justitia: A Comprehensive Survey on Safety Evaluation of LLMs LLM-PBE: Assessing Data Privacy in Large Language Models
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 01061d32-ffd0-4517-a929-7490bfdff1c3 · outbound
The Scales of Justitia: A Comprehensive Survey on Safety Evaluation of LLMs MedSafetyBench: Evaluating and Improving the Medical Safety of Large Language Models
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 78188daa-cfd8-411e-b556-0bf09ef0c4f5 · outbound
The Scales of Justitia: A Comprehensive Survey on Safety Evaluation of LLMs Purple Llama CyberSecEval: A Secure Coding Benchmark for Language Models
Reference 57
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a5c71bf5-c1c4-4f94-841b-1543d665cf28 · outbound
The Scales of Justitia: A Comprehensive Survey on Safety Evaluation of LLMs Cweval: Outcome- driven evaluation on functionality and security of llm code generation,
Reference 58
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 61f9837f-bff1-4280-b78e-f6453c9785b0 · outbound
The Scales of Justitia: A Comprehensive Survey on Safety Evaluation of LLMs SafeGenBench: A Benchmark Framework for Security Vulnerability Detection in LLM-Generated Code
Reference 59
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation af3df8c1-40a0-46ea-be95-24be8b49e5bd · outbound
The Scales of Justitia: A Comprehensive Survey on Safety Evaluation of LLMs SHIELD: Evaluation and Defense Strategies for Copyright Compliance in LLM Text Generation
Reference 60
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0dda50f5-b5d6-496b-9d01-6bbc18b3d89b · outbound
The Scales of Justitia: A Comprehensive Survey on Safety Evaluation of LLMs TRIDENT: Benchmarking LLM Safety in Finance, Medicine, and Law
Reference 61
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 523e7bf9-488c-4f03-955d-551ed5e6783d · outbound
The Scales of Justitia: A Comprehensive Survey on Safety Evaluation of LLMs Rethinking How to Evaluate Language Model Jailbreak
Reference 62
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 02c0fdab-1e14-4b37-9335-3bfda1829338 · outbound
The Scales of Justitia: A Comprehensive Survey on Safety Evaluation of LLMs A StrongREJECT for Empty Jailbreaks
Reference 63
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6a1bdefb-1b64-4581-99ed-cd19080fb6c1 · outbound
The Scales of Justitia: A Comprehensive Survey on Safety Evaluation of LLMs "Not Aligned" is Not "Malicious": Being Careful about Hallucinations of Large Language Models' Jailbreak
Reference 64
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 266333ba-7c41-48f7-85d0-e27a04f77d93 · outbound
The Scales of Justitia: A Comprehensive Survey on Safety Evaluation of LLMs JailbreakRadar: Comprehensive Assessment of Jailbreak Attacks Against LLMs
Reference 65
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c9e94feb-06c5-456c-9bf8-bdcfb066834b · outbound
The Scales of Justitia: A Comprehensive Survey on Safety Evaluation of LLMs R-Judge: Benchmarking Safety Risk Awareness for LLM Agents
Reference 66
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 18774a35-b8c8-4437-ab17-31ec03d9fba4 · outbound
The Scales of Justitia: A Comprehensive Survey on Safety Evaluation of LLMs Do-not-answer: Evaluating safeguards in llms,
Reference 67
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 57447a39-ddb6-4026-9b68-e17685ee28aa · outbound
The Scales of Justitia: A Comprehensive Survey on Safety Evaluation of LLMs CoSafe: Evaluating Large Language Model Safety in Multi-Turn Dialogue Coreference
Reference 68
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 39f87ec4-b572-468a-8ca6-49826f53a9bd · outbound
The Scales of Justitia: A Comprehensive Survey on Safety Evaluation of LLMs JAILJUDGE: A Comprehensive Jailbreak Judge Benchmark with Multi-Agent Enhanced Explanation Evaluation Framework
Reference 69
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5d669970-5c06-4d74-98f9-d4fed5a5a342 · outbound
The Scales of Justitia: A Comprehensive Survey on Safety Evaluation of LLMs SORRY-Bench: Systematically Evaluating Large Language Model Safety Refusal
Reference 70
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ce2e4144-2beb-4ad1-a941-531422c683b6 · outbound
The Scales of Justitia: A Comprehensive Survey on Safety Evaluation of LLMs AEGIS: Online Adaptive AI Content Safety Moderation with Ensemble of LLM Experts
Reference 71
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d7eb6bf1-2690-4630-8596-899c53113dab · outbound
The Scales of Justitia: A Comprehensive Survey on Safety Evaluation of LLMs SALAD-Bench: A Hierarchical and Comprehensive Safety Benchmark for Large Language Models
Reference 72
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 572e6694-eddc-49d5-851b-3534c1deaa90 · outbound
The Scales of Justitia: A Comprehensive Survey on Safety Evaluation of LLMs ALERT: A Comprehensive Benchmark for Assessing Large Language Models' Safety through Red Teaming
Reference 73
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cf1de549-be79-4d74-a134-3f4e255ace61 · outbound
The Scales of Justitia: A Comprehensive Survey on Safety Evaluation of LLMs WildGuard: Open One-Stop Moderation Tools for Safety Risks, Jailbreaks, and Refusals of LLMs
Reference 74
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8ba49cd3-3476-4039-9bd9-4f94ab3c2c62 · outbound
The Scales of Justitia: A Comprehensive Survey on Safety Evaluation of LLMs "Do Anything Now": Characterizing and Evaluating In-The-Wild Jailbreak Prompts on Large Language Models
Reference 75
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation add6c427-c3f2-48e9-81f7-29d979303741 · outbound
The Scales of Justitia: A Comprehensive Survey on Safety Evaluation of LLMs CHiSafetyBench: A Chinese Hierarchical Safety Benchmark for Large Language Models
Reference 76
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5b538f84-aa5b-4c16-8f9c-b01e0a256f96 · outbound
The Scales of Justitia: A Comprehensive Survey on Safety Evaluation of LLMs A Chinese Dataset for Evaluating the Safeguards in Large Language Models
Reference 77
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5e22d829-7044-4107-85cd-2be2c5548390 · outbound
The Scales of Justitia: A Comprehensive Survey on Safety Evaluation of LLMs JailBench: A Comprehensive Chinese Security Assessment Benchmark for Large Language Models
Reference 78
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ca731799-ea27-4e7a-9c7b-1f1e839cc98a · outbound
The Scales of Justitia: A Comprehensive Survey on Safety Evaluation of LLMs CRiskEval: A Chinese Multi-Level Risk Evaluation Benchmark Dataset for Large Language Models
Reference 79
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5ff719d5-d45f-4af8-8b86-829aab8d9f95 · outbound
The Scales of Justitia: A Comprehensive Survey on Safety Evaluation of LLMs ChineseSafe: A Chinese Benchmark for Evaluating Safety in Large Language Models
Reference 80
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4d02dc49-52e6-4edc-bd31-c0d406093cc2 · outbound
The Scales of Justitia: A Comprehensive Survey on Safety Evaluation of LLMs Latent Jailbreak: A Benchmark for Evaluating Text Safety and Output Robustness of Large Language Models
Reference 81
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b06fc21c-6af4-4cd9-8846-993cfd87ade3 · outbound
The Scales of Justitia: A Comprehensive Survey on Safety Evaluation of LLMs Safetybench: Evaluating the safety of large language models,
Reference 82
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d4ac6a66-df10-4627-9818-94933ba03a0e · outbound
The Scales of Justitia: A Comprehensive Survey on Safety Evaluation of LLMs S-Eval: Towards Automated and Comprehensive Safety Evaluation for Large Language Models
Reference 83
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ba7ea828-025e-419d-8261-7f473360d64b · outbound
The Scales of Justitia: A Comprehensive Survey on Safety Evaluation of LLMs AraTrust: An Evaluation of Trustworthiness for LLMs in Arabic
Reference 84
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5cf06d1a-4497-4622-81af-3c946d84e8ef · outbound
The Scales of Justitia: A Comprehensive Survey on Safety Evaluation of LLMs AnswerCarefully: A Dataset for Improving the Safety of Japanese LLM Output
Reference 85
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a560036a-b759-40a8-94fa-e2463d811690 · outbound
The Scales of Justitia: A Comprehensive Survey on Safety Evaluation of LLMs All languages matter: On the multilingual safety of llms,
Reference 86
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0a192469-b02d-4b1a-8cf1-e5ce6f13696e · outbound
The Scales of Justitia: A Comprehensive Survey on Safety Evaluation of LLMs Annotation alignment: Comparing LLM and human annotations of conversational safety
Reference 87
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d69db30b-4bb4-4c01-aa10-698b3e24da60 · outbound
The Scales of Justitia: A Comprehensive Survey on Safety Evaluation of LLMs Making them ask and answer: Jailbreaking large language models in few queries via disguise and reconstruction,
Reference 88
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 20c6fe58-3282-48ed-ad79-6b3d4083e531 · outbound
The Scales of Justitia: A Comprehensive Survey on Safety Evaluation of LLMs AutoDefense: Multi-Agent LLM Defense against Jailbreak Attacks
Reference 89
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 59200a37-2356-4583-acc2-ca972c923fea · outbound
The Scales of Justitia: A Comprehensive Survey on Safety Evaluation of LLMs OpenAI Moderation Endpoint,
Reference 90
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7a5329e4-ad5b-4d08-aa82-9428ead862d0 · outbound
The Scales of Justitia: A Comprehensive Survey on Safety Evaluation of LLMs Google Perspective API,
Reference 91
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 75da5195-c3f9-48c7-8bf0-0314bf98f532 · outbound
The Scales of Justitia: A Comprehensive Survey on Safety Evaluation of LLMs Microsoft Azure AI Content Safety API,
Reference 92
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ac0c4c18-e839-48df-8d70-87d8e6188ef2 · outbound
The Scales of Justitia: A Comprehensive Survey on Safety Evaluation of LLMs GPTFUZZER: Red Teaming Large Language Models with Auto-Generated Jailbreak Prompts
Reference 93
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e18b27fe-1e54-491a-a7c3-874bbcf659ff · outbound
The Scales of Justitia: A Comprehensive Survey on Safety Evaluation of LLMs Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!
Reference 94
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6627fce0-2607-4a62-b456-df324fc6b2e2 · outbound
The Scales of Justitia: A Comprehensive Survey on Safety Evaluation of LLMs Jailbreaking Black Box Large Language Models in Twenty Queries
Reference 95
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3c351043-6e22-4ba2-8063-e415567b38df · outbound
The Scales of Justitia: A Comprehensive Survey on Safety Evaluation of LLMs Scalable and Transferable Black-Box Jailbreaks for Language Models via Persona Modulation
Reference 96
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 424edd04-e787-49a4-8572-4b0d520bdb7d · outbound
The Scales of Justitia: A Comprehensive Survey on Safety Evaluation of LLMs Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations
Reference 97
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fb742edd-25b6-40c3-8c08-6af92546d598 · outbound
The Scales of Justitia: A Comprehensive Survey on Safety Evaluation of LLMs Meta llama guard 2,
Reference 98
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cf752b6e-8bba-46ef-8dac-e08e997d178e · outbound
The Scales of Justitia: A Comprehensive Survey on Safety Evaluation of LLMs The llama 3 family of models,
Reference 99
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d80a1f43-2742-44ff-a2a1-31cf88ef4dc3 · outbound
The Scales of Justitia: A Comprehensive Survey on Safety Evaluation of LLMs The Llama 3 Herd of Models
Reference 100
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 02d67783-e274-46fb-8f3c-c6728334936a · inbound
PRISON: Unmasking the Criminal Potential of Large Language Models The Scales of Justitia: A Comprehensive Survey on Safety Evaluation of LLMs
Reference 2025
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 36cd4c5f-04c3-46b8-9ee2-b89c5edef328 · inbound
LLMEval-Fair: A Large-Scale Longitudinal Study on Robust and Fair Evaluation of Large Language Models The Scales of Justitia: A Comprehensive Survey on Safety Evaluation of LLMs
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation be0be778-36e8-4acb-b494-9b96d9cca347 · inbound
Safe for Whom? Rethinking How We Evaluate the Safety of LLMs for Real Users The Scales of Justitia: A Comprehensive Survey on Safety Evaluation of LLMs
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation b5b39288-6849-4912-bf6c-b08044cad947 · inbound
When Safety Becomes a Vulnerability: Exploiting LLM Alignment Homogeneity for Transferable Blocking in RAG The Scales of Justitia: A Comprehensive Survey on Safety Evaluation of LLMs
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3ec71ee7-047f-488c-bf7a-63988fe1b47c · inbound
Steering at the Source: Style Modulation Heads for Robust Persona Control The Scales of Justitia: A Comprehensive Survey on Safety Evaluation of LLMs
Reference 2023
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0a90883e-cb2f-478e-b566-4c22d751e037 · inbound
A Systematic Evaluation of Black-Box Uncertainty Estimation Methods for Large Language Models The Scales of Justitia: A Comprehensive Survey on Safety Evaluation of LLMs
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 23fe07c8-e95d-47bc-adea-783bcbb938f0 · inbound
Efficient Safety Benchmarking via Item Response Theory The Scales of Justitia: A Comprehensive Survey on Safety Evaluation of LLMs
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 0c4dbdd8-cfb3-482e-9fd6-5dcf00b4372c · inbound
Testing Retrieval-Augmented Generation Systems with Chunk Coverage The Scales of Justitia: A Comprehensive Survey on Safety Evaluation of LLMs
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d46ec9ac-96fe-416b-9832-b160189c2997 · inbound
Large Language Models and Social Media Information Integrity: Opportunities, Challenges, and Research Directions The Scales of Justitia: A Comprehensive Survey on Safety Evaluation of LLMs
Reference 112
Source-reported events for the cited work
Unavailable: canonical work link unavailable.