Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T14:52:31.262024Z
Paper Citation Record · LEDGER
As of 8 August 2026, this Paper Citation Record lists 58 of 58 outbound references and 0 inbound Pith citation observations for arXiv:2505.17332.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T14:52:31.262024Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
58 of 58 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 26b1aae7-672a-45e4-b9c2-5039e9325947 · outbound
SweEval: Do LLMs Really Swear? A Safety Benchmark for Testing Limits for Enterprise Use online" 'onlinestring :=
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7c087a5a-b0ad-4bcf-b766-8d2e8440c8da · outbound
SweEval: Do LLMs Really Swear? A Safety Benchmark for Testing Limits for Enterprise Use write newline
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 64a2ae5a-527d-47ca-b002-dc740e031058 · outbound
SweEval: Do LLMs Really Swear? A Safety Benchmark for Testing Limits for Enterprise Use Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c2631a25-09ad-4c04-a86b-ff246c00d569 · outbound
SweEval: Do LLMs Really Swear? A Safety Benchmark for Testing Limits for Enterprise Use Unresolved cited work
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation d0965ad8-1162-4095-8b4a-402b8dc9755b · outbound
SweEval: Do LLMs Really Swear? A Safety Benchmark for Testing Limits for Enterprise Use MVTamperBench: Evaluating Robustness of Vision-Language Models
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e895a1df-56ae-4bd4-a19c-9c5d42c6c30f · outbound
SweEval: Do LLMs Really Swear? A Safety Benchmark for Testing Limits for Enterprise Use Unresolved cited work
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation eb97f4e2-8ef0-4ace-883d-09589e055212 · outbound
SweEval: Do LLMs Really Swear? A Safety Benchmark for Testing Limits for Enterprise Use Unresolved cited work
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 550c9ee5-1b49-4834-93b7-21747f8a4949 · outbound
SweEval: Do LLMs Really Swear? A Safety Benchmark for Testing Limits for Enterprise Use Unresolved cited work
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6ee05890-3ee7-49f1-ac40-767856edd6ea · outbound
SweEval: Do LLMs Really Swear? A Safety Benchmark for Testing Limits for Enterprise Use Enhancing Document AI Data Generation Through Graph-Based Synthetic Layouts
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 207e7e05-166f-44cc-b4a5-d9f29d35b8d7 · outbound
SweEval: Do LLMs Really Swear? A Safety Benchmark for Testing Limits for Enterprise Use Do, Yan Xu, and Pascale Fung
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 30a523f3-53cd-48df-9482-3c14107d1c63 · outbound
SweEval: Do LLMs Really Swear? A Safety Benchmark for Testing Limits for Enterprise Use Language Models are Few-Shot Learners
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a56ae450-2731-4a80-a54a-990f265db275 · outbound
SweEval: Do LLMs Really Swear? A Safety Benchmark for Testing Limits for Enterprise Use JailbreakBench: An Open Robustness Benchmark for Jailbreaking Large Language Models
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f7d9ba58-cc46-45fd-a320-59388953a377 · outbound
SweEval: Do LLMs Really Swear? A Safety Benchmark for Testing Limits for Enterprise Use Crosslingual Capabilities and Knowledge Barriers in Multilingual Large Language Models
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7f04b227-b7e4-436e-b7d4-918c16af380e · outbound
SweEval: Do LLMs Really Swear? A Safety Benchmark for Testing Limits for Enterprise Use Unresolved cited work
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a93f0125-3504-4efb-a012-be1992f6f87d · outbound
SweEval: Do LLMs Really Swear? A Safety Benchmark for Testing Limits for Enterprise Use The Llama 3 Herd of Models
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4354f0a9-961f-4be5-b27d-78afba2db335 · outbound
SweEval: Do LLMs Really Swear? A Safety Benchmark for Testing Limits for Enterprise Use Multilingual Large Language Models and Curse of Multilinguality
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 556dabd3-900e-4d59-a3f7-865468e7a661 · outbound
SweEval: Do LLMs Really Swear? A Safety Benchmark for Testing Limits for Enterprise Use ToxiGen: A Large-Scale Machine-Generated Dataset for Adversarial and Implicit Hate Speech Detection
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fe891c49-9c68-4b27-a31d-636d3040cd64 · outbound
SweEval: Do LLMs Really Swear? A Safety Benchmark for Testing Limits for Enterprise Use PKU-SafeRLHF: Towards Multi-Level Safety Alignment for LLMs with Human Preference
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 289d9c3d-4374-49f1-90c9-abe2b0413ec1 · outbound
SweEval: Do LLMs Really Swear? A Safety Benchmark for Testing Limits for Enterprise Use Mistral 7B
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e14fc6fb-c632-433d-b421-a9306e3bc01b · outbound
SweEval: Do LLMs Really Swear? A Safety Benchmark for Testing Limits for Enterprise Use A Survey on Large Language Models for Code Generation
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bb0d6a1a-e472-4331-a9bb-dd6f620818a6 · outbound
SweEval: Do LLMs Really Swear? A Safety Benchmark for Testing Limits for Enterprise Use Unresolved cited work
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d95e5fad-f1cd-4af4-9819-486e74c05bc6 · outbound
SweEval: Do LLMs Really Swear? A Safety Benchmark for Testing Limits for Enterprise Use Fine-Tuning, Quantization, and LLMs: Navigating Unintended Outcomes
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7e09b11c-0844-42b0-9106-664da1d3e6ec · outbound
SweEval: Do LLMs Really Swear? A Safety Benchmark for Testing Limits for Enterprise Use SALAD-Bench: A Hierarchical and Comprehensive Safety Benchmark for Large Language Models
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3ffc685f-ab4a-4baa-923c-52ef5ab7142b · outbound
SweEval: Do LLMs Really Swear? A Safety Benchmark for Testing Limits for Enterprise Use Controllable Text Generation for Large Language Models: A Survey
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e773cb32-b65a-4db0-8695-1cd4132b6e08 · outbound
SweEval: Do LLMs Really Swear? A Safety Benchmark for Testing Limits for Enterprise Use ToxicChat: Unveiling Hidden Challenges of Toxicity Detection in Real-World User-AI Conversation
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8ada4d7c-2f56-49aa-8e40-16cf10badd7a · outbound
SweEval: Do LLMs Really Swear? A Safety Benchmark for Testing Limits for Enterprise Use Corporate Communication Companion (CCC): An LLM-empowered Writing Assistant for Workplace Social Media
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation da7635fc-2190-4561-b9c3-f2daa71a078a · outbound
SweEval: Do LLMs Really Swear? A Safety Benchmark for Testing Limits for Enterprise Use A Paradigm Shift: The Future of Machine Translation Lies with Large Language Models
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e70cb713-e48b-4f6a-8c0b-d533273190ac · outbound
SweEval: Do LLMs Really Swear? A Safety Benchmark for Testing Limits for Enterprise Use HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust Refusal
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 96c1bc5c-be2d-424d-b62f-1c8a15e68482 · outbound
SweEval: Do LLMs Really Swear? A Safety Benchmark for Testing Limits for Enterprise Use Unresolved cited work
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation daf4a831-a6f4-4279-b27f-ce3504892117 · outbound
SweEval: Do LLMs Really Swear? A Safety Benchmark for Testing Limits for Enterprise Use Unresolved cited work
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation af3517cd-6827-4fff-a2d1-0ce3ba1f690d · outbound
SweEval: Do LLMs Really Swear? A Safety Benchmark for Testing Limits for Enterprise Use Unresolved cited work
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b52a48f3-7d90-42a4-a53f-80ae4217600b · outbound
SweEval: Do LLMs Really Swear? A Safety Benchmark for Testing Limits for Enterprise Use LLM for Barcodes: Generating Diverse Synthetic Data for Identity Documents
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fa05beb8-01b6-4ffc-96c4-a948895bb029 · outbound
SweEval: Do LLMs Really Swear? A Safety Benchmark for Testing Limits for Enterprise Use Review of reference generation methods in large language models
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7090315f-e05e-42f2-8158-447c8d1cf6fd · outbound
SweEval: Do LLMs Really Swear? A Safety Benchmark for Testing Limits for Enterprise Use Unresolved cited work
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5b2fa618-4552-4e31-8336-e1f9e5468b8e · outbound
SweEval: Do LLMs Really Swear? A Safety Benchmark for Testing Limits for Enterprise Use Tokenization Matters: Improving Zero-Shot NER for Indic Languages
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 686ffdf8-d601-414b-ac05-0c6020e20bd6 · outbound
SweEval: Do LLMs Really Swear? A Safety Benchmark for Testing Limits for Enterprise Use Clinical QA 2.0: Multi-Task Learning for Answer Extraction and Categorization
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c9cdb51a-fff7-4bf0-8352-94db76df3dcc · outbound
SweEval: Do LLMs Really Swear? A Safety Benchmark for Testing Limits for Enterprise Use Survey of Large Multimodal Model Datasets, Application Categories and Taxonomy
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 69935a2f-d505-4a50-b4ac-003d4e18c356 · outbound
SweEval: Do LLMs Really Swear? A Safety Benchmark for Testing Limits for Enterprise Use Unresolved cited work
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 52b6447c-c493-4b19-89bc-6796518e7df5 · outbound
SweEval: Do LLMs Really Swear? A Safety Benchmark for Testing Limits for Enterprise Use XSTest: A Test Suite for Identifying Exaggerated Safety Behaviours in Large Language Models
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5058c140-2e80-4716-b643-cb7382fe309d · outbound
SweEval: Do LLMs Really Swear? A Safety Benchmark for Testing Limits for Enterprise Use Unresolved cited work
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2b1e54c6-b92e-48a4-aefe-4140254a9bb5 · outbound
SweEval: Do LLMs Really Swear? A Safety Benchmark for Testing Limits for Enterprise Use The Language Barrier: Dissecting Safety Challenges of LLMs in Multilingual Contexts
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4a3b8bd0-cde7-4ee9-a71d-07972626824b · outbound
SweEval: Do LLMs Really Swear? A Safety Benchmark for Testing Limits for Enterprise Use Unresolved cited work
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 6875d0ee-2f5a-41e4-a4c4-65e207abcc1f · outbound
SweEval: Do LLMs Really Swear? A Safety Benchmark for Testing Limits for Enterprise Use Text Classification via Large Language Models
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7e4a57dd-0d8c-41dd-b295-f35e2e2bdbe4 · outbound
SweEval: Do LLMs Really Swear? A Safety Benchmark for Testing Limits for Enterprise Use Unresolved cited work
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9ad2ab34-9314-4d9f-966e-fba96dd00264 · outbound
SweEval: Do LLMs Really Swear? A Safety Benchmark for Testing Limits for Enterprise Use ALERT: A Comprehensive Benchmark for Assessing Large Language Models' Safety through Red Teaming
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fff2e369-c818-4d8d-b4a1-3064e957e2bb · outbound
SweEval: Do LLMs Really Swear? A Safety Benchmark for Testing Limits for Enterprise Use Unresolved cited work
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 8f548a86-a549-49d8-bf54-2da948643e91 · outbound
SweEval: Do LLMs Really Swear? A Safety Benchmark for Testing Limits for Enterprise Use LLaMA: Open and Efficient Foundation Language Models
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5661c39f-23e3-44c4-8c81-356c52133435 · outbound
SweEval: Do LLMs Really Swear? A Safety Benchmark for Testing Limits for Enterprise Use All Languages Matter: On the Multilingual Safety of Large Language Models
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8dea832e-a657-4a75-9287-1323bcb16dd5 · outbound
SweEval: Do LLMs Really Swear? A Safety Benchmark for Testing Limits for Enterprise Use Do-Not-Answer: A Dataset for Evaluating Safeguards in LLMs
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a4d6c8c2-935d-472d-a3e9-0023f6223a01 · outbound
SweEval: Do LLMs Really Swear? A Safety Benchmark for Testing Limits for Enterprise Use Adaptable and Reliable Text Classification using Large Language Models
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c99df6db-c545-4525-8675-69f32fd72396 · outbound
SweEval: Do LLMs Really Swear? A Safety Benchmark for Testing Limits for Enterprise Use Unresolved cited work
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation fa468d64-aa82-41e4-93b2-5c3567b007d3 · outbound
SweEval: Do LLMs Really Swear? A Safety Benchmark for Testing Limits for Enterprise Use Unresolved cited work
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 6d811496-4a3b-4978-864f-92c658110c44 · outbound
SweEval: Do LLMs Really Swear? A Safety Benchmark for Testing Limits for Enterprise Use SORRY-Bench: Systematically Evaluating Large Language Model Safety Refusal
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4b32d5de-e91b-44ca-bec3-3bf79eb58ad6 · outbound
SweEval: Do LLMs Really Swear? A Safety Benchmark for Testing Limits for Enterprise Use Unresolved cited work
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6df9c3e1-659f-4b29-9f90-bbf5a8fe15e0 · outbound
SweEval: Do LLMs Really Swear? A Safety Benchmark for Testing Limits for Enterprise Use Unresolved cited work
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a444735e-47da-4991-9cee-16ece478ec49 · outbound
SweEval: Do LLMs Really Swear? A Safety Benchmark for Testing Limits for Enterprise Use SafetyBench: Evaluating the Safety of Large Language Models
Reference 57
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a1c9744b-1e2d-4831-be5f-977c415f43c2 · outbound
SweEval: Do LLMs Really Swear? A Safety Benchmark for Testing Limits for Enterprise Use Unresolved cited work
Reference 58
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 419e77c1-3e29-460a-b529-1c910287c98e · outbound
SweEval: Do LLMs Really Swear? A Safety Benchmark for Testing Limits for Enterprise Use Universal and Transferable Adversarial Attacks on Aligned Language Models
Reference 59
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.