Pith. sign in

Paper Citation Record · LEDGER

SweEval: Do LLMs Really Swear? A Safety Benchmark for Testing Limits for Enterprise Use

As of 8 August 2026, this Paper Citation Record lists 58 of 58 outbound references and 0 inbound Pith citation observations for arXiv:2505.17332.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.17332 v1

Coverage vector

measured 58 of 58 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:52:31.262024Z

measured 58 of 58 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

58 of 58 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved58
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 26b1aae7-672a-45e4-b9c2-5039e9325947 · outbound

This paper cites online" 'onlinestring :=.

SweEval: Do LLMs Really Swear? A Safety Benchmark for Testing Limits for Enterprise Use online" 'onlinestring :=

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T14:52:26.322302Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:52:26.322302Z digest=sha256:5b9775345fe901f27890031fa09c0f545a0a11c7f286fb72ae2f600f7b2242b8

Observation 7c087a5a-b0ad-4bcf-b766-8d2e8440c8da · outbound

This paper cites write newline.

SweEval: Do LLMs Really Swear? A Safety Benchmark for Testing Limits for Enterprise Use write newline

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T14:52:26.464178Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:52:26.464178Z digest=sha256:1b30b0be14316addaa214a223614780e7a81eec1adb67c5598eabfbb640f5845

Observation 64a2ae5a-527d-47ca-b002-dc740e031058 · outbound

This paper cites Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone.

SweEval: Do LLMs Really Swear? A Safety Benchmark for Testing Limits for Enterprise Use Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T14:52:26.615041Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:52:26.615041Z digest=sha256:cde200ff319d9552295b0bd26f59fe062937b2fd67eadbff55894cad264d205d

Observation c2631a25-09ad-4c04-a86b-ff246c00d569 · outbound

This paper cites an unresolved cited work.

SweEval: Do LLMs Really Swear? A Safety Benchmark for Testing Limits for Enterprise Use Unresolved cited work

Reference 4

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:52:33.061472Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T14:52:26.750767Z digest=sha256:cfb73c2355a54618d32dc67dfb554c18712f12fcd6ce3229691c7ea317cb4ab0

Observation d0965ad8-1162-4095-8b4a-402b8dc9755b · outbound

This paper cites MVTamperBench: Evaluating Robustness of Vision-Language Models.

SweEval: Do LLMs Really Swear? A Safety Benchmark for Testing Limits for Enterprise Use MVTamperBench: Evaluating Robustness of Vision-Language Models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T14:52:26.868304Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:52:26.868304Z digest=sha256:2828fba62db11da698dd5494d24c64e75b96a6dae439a46ebc6246e2c0f84fab

Observation e895a1df-56ae-4bd4-a19c-9c5d42c6c30f · outbound

This paper cites an unresolved cited work.

SweEval: Do LLMs Really Swear? A Safety Benchmark for Testing Limits for Enterprise Use Unresolved cited work

Reference 6

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:52:32.928539Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T14:52:26.976950Z digest=sha256:ed44bbb275fc77c70ac31834f52fe6ec5fdcf743e4f9e410463bb6b98cbdea63

Observation eb97f4e2-8ef0-4ace-883d-09589e055212 · outbound

This paper cites an unresolved cited work.

SweEval: Do LLMs Really Swear? A Safety Benchmark for Testing Limits for Enterprise Use Unresolved cited work

Reference 7

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:52:32.801949Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T14:52:27.115403Z digest=sha256:5634480585e6cb371ee45d339d73109f1c69f3bea6c12c22fa6648f645126c89

Observation 550c9ee5-1b49-4834-93b7-21747f8a4949 · outbound

This paper cites an unresolved cited work.

SweEval: Do LLMs Really Swear? A Safety Benchmark for Testing Limits for Enterprise Use Unresolved cited work

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T14:52:27.165158Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:52:27.165158Z digest=sha256:6260f2766df252568a05cb1e709e9c65123fc011a74fec3bf32ca6c29f7173f4

Observation 6ee05890-3ee7-49f1-ac40-767856edd6ea · outbound

This paper cites Enhancing Document AI Data Generation Through Graph-Based Synthetic Layouts.

SweEval: Do LLMs Really Swear? A Safety Benchmark for Testing Limits for Enterprise Use Enhancing Document AI Data Generation Through Graph-Based Synthetic Layouts

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T14:52:27.217122Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:52:27.217122Z digest=sha256:aa0e6d343bc40508f9b4b762bd3a5e69e1feda7dbf15d637196627682e486527

Observation 207e7e05-166f-44cc-b4a5-d9f29d35b8d7 · outbound

This paper cites Do, Yan Xu, and Pascale Fung.

SweEval: Do LLMs Really Swear? A Safety Benchmark for Testing Limits for Enterprise Use Do, Yan Xu, and Pascale Fung

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T14:52:27.279376Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:52:27.279376Z digest=sha256:c1b2170557e45c36c584c18eec6e7e6dc1c318ba7256d279392a8700d366d360

Observation 30a523f3-53cd-48df-9482-3c14107d1c63 · outbound

This paper cites Language Models are Few-Shot Learners.

SweEval: Do LLMs Really Swear? A Safety Benchmark for Testing Limits for Enterprise Use Language Models are Few-Shot Learners

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T14:52:27.438818Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:52:27.438818Z digest=sha256:44de4c4f983852c8628f609f1c88c8fcf9ed67c24f0cca206768d3a651d17790

Observation a56ae450-2731-4a80-a54a-990f265db275 · outbound

This paper cites JailbreakBench: An Open Robustness Benchmark for Jailbreaking Large Language Models.

SweEval: Do LLMs Really Swear? A Safety Benchmark for Testing Limits for Enterprise Use JailbreakBench: An Open Robustness Benchmark for Jailbreaking Large Language Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T14:52:27.510965Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:52:27.510965Z digest=sha256:881332ce731a2e9aecc57d77e70af3e7cf115b7cf81018918f9dfe339639f89a

Observation f7d9ba58-cc46-45fd-a320-59388953a377 · outbound

This paper cites Crosslingual Capabilities and Knowledge Barriers in Multilingual Large Language Models.

SweEval: Do LLMs Really Swear? A Safety Benchmark for Testing Limits for Enterprise Use Crosslingual Capabilities and Knowledge Barriers in Multilingual Large Language Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T14:52:27.565303Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:52:27.565303Z digest=sha256:9a8d092feed284718f92c48780cd2e6f1556625996d80f59a6a86ab39828819d

Observation 7f04b227-b7e4-436e-b7d4-918c16af380e · outbound

This paper cites an unresolved cited work.

SweEval: Do LLMs Really Swear? A Safety Benchmark for Testing Limits for Enterprise Use Unresolved cited work

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T14:52:27.646088Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:52:27.646088Z digest=sha256:9e2abc613fe0b0f7fb8eba20af570cfaea46a4e1ccbac3534855c88c16b85aa8

Observation a93f0125-3504-4efb-a012-be1992f6f87d · outbound

This paper cites The Llama 3 Herd of Models.

SweEval: Do LLMs Really Swear? A Safety Benchmark for Testing Limits for Enterprise Use The Llama 3 Herd of Models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T14:52:27.742920Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:52:27.742920Z digest=sha256:2e8859287d8a3552311b402c3ccc98c298c5e341b7092daad8dc728e5b921f0e

Observation 4354f0a9-961f-4be5-b27d-78afba2db335 · outbound

This paper cites Multilingual Large Language Models and Curse of Multilinguality.

SweEval: Do LLMs Really Swear? A Safety Benchmark for Testing Limits for Enterprise Use Multilingual Large Language Models and Curse of Multilinguality

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T14:52:27.834797Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:52:27.834797Z digest=sha256:d7e8fd7f22486d9033175bde5db15006b3c90998414022b5113d3b5a5ed8b031

Observation 556dabd3-900e-4d59-a3f7-865468e7a661 · outbound

This paper cites ToxiGen: A Large-Scale Machine-Generated Dataset for Adversarial and Implicit Hate Speech Detection.

SweEval: Do LLMs Really Swear? A Safety Benchmark for Testing Limits for Enterprise Use ToxiGen: A Large-Scale Machine-Generated Dataset for Adversarial and Implicit Hate Speech Detection

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T14:52:27.948003Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:52:27.948003Z digest=sha256:0036ed2858f6a4f4d91d7bfcb0682fe7ef691a60992f37d178b22a79f66ce8ea

Observation fe891c49-9c68-4b27-a31d-636d3040cd64 · outbound

This paper cites PKU-SafeRLHF: Towards Multi-Level Safety Alignment for LLMs with Human Preference.

SweEval: Do LLMs Really Swear? A Safety Benchmark for Testing Limits for Enterprise Use PKU-SafeRLHF: Towards Multi-Level Safety Alignment for LLMs with Human Preference

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T14:52:28.092591Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:52:28.092591Z digest=sha256:19df09de2b986d0d3e7d03166ee6aab148481800bd1283446d617943b33b07e4

Observation 289d9c3d-4374-49f1-90c9-abe2b0413ec1 · outbound

This paper cites Mistral 7B.

SweEval: Do LLMs Really Swear? A Safety Benchmark for Testing Limits for Enterprise Use Mistral 7B

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T14:52:28.202392Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:52:28.202392Z digest=sha256:e52ccbbb2b75a97d0a474a861fdfeb62d29744b9de6d72950b64555f6b07ff4b

Observation e14fc6fb-c632-433d-b421-a9306e3bc01b · outbound

This paper cites A Survey on Large Language Models for Code Generation.

SweEval: Do LLMs Really Swear? A Safety Benchmark for Testing Limits for Enterprise Use A Survey on Large Language Models for Code Generation

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T14:52:28.304469Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:52:28.304469Z digest=sha256:612874ffe26e280e6ea8ea789d3e6265e91626f1c441bc81702de8fdef68a663

Observation bb0d6a1a-e472-4331-a9bb-dd6f620818a6 · outbound

This paper cites an unresolved cited work.

SweEval: Do LLMs Really Swear? A Safety Benchmark for Testing Limits for Enterprise Use Unresolved cited work

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T14:52:28.406960Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:52:28.406960Z digest=sha256:685d2898541d1efe5613a45e38b4d3b4ed1a1bac6dee05156d4d99c153ed65d5

Observation d95e5fad-f1cd-4af4-9819-486e74c05bc6 · outbound

This paper cites Fine-Tuning, Quantization, and LLMs: Navigating Unintended Outcomes.

SweEval: Do LLMs Really Swear? A Safety Benchmark for Testing Limits for Enterprise Use Fine-Tuning, Quantization, and LLMs: Navigating Unintended Outcomes

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T14:52:28.533469Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:52:28.533469Z digest=sha256:75c9f8104dc293ee39523d64b6f61e36fe961d237305be9c5a7f784a01d80d19

Observation 7e09b11c-0844-42b0-9106-664da1d3e6ec · outbound

This paper cites SALAD-Bench: A Hierarchical and Comprehensive Safety Benchmark for Large Language Models.

SweEval: Do LLMs Really Swear? A Safety Benchmark for Testing Limits for Enterprise Use SALAD-Bench: A Hierarchical and Comprehensive Safety Benchmark for Large Language Models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T14:52:28.676227Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:52:28.676227Z digest=sha256:06e8ada90273cfd60bce88cc2054fdf9acc19e1426662d60b4df9f04b788ef56

Observation 3ffc685f-ab4a-4baa-923c-52ef5ab7142b · outbound

This paper cites Controllable Text Generation for Large Language Models: A Survey.

SweEval: Do LLMs Really Swear? A Safety Benchmark for Testing Limits for Enterprise Use Controllable Text Generation for Large Language Models: A Survey

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T14:52:28.795914Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:52:28.795914Z digest=sha256:18723231fb428e7a658e8e566fee0e21f558f9c521b7febe87727f783a9d7ab2

Observation e773cb32-b65a-4db0-8695-1cd4132b6e08 · outbound

This paper cites ToxicChat: Unveiling Hidden Challenges of Toxicity Detection in Real-World User-AI Conversation.

SweEval: Do LLMs Really Swear? A Safety Benchmark for Testing Limits for Enterprise Use ToxicChat: Unveiling Hidden Challenges of Toxicity Detection in Real-World User-AI Conversation

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T14:52:28.909375Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:52:28.909375Z digest=sha256:82036c0af1fa3c865373cd86edc577e533465fdaa308c3a67c82014d9fdf63ab

Observation 8ada4d7c-2f56-49aa-8e40-16cf10badd7a · outbound

This paper cites Corporate Communication Companion (CCC): An LLM-empowered Writing Assistant for Workplace Social Media.

SweEval: Do LLMs Really Swear? A Safety Benchmark for Testing Limits for Enterprise Use Corporate Communication Companion (CCC): An LLM-empowered Writing Assistant for Workplace Social Media

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T14:52:28.995452Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:52:28.995452Z digest=sha256:55e95e778b3f90428a958ed06c8cd40929addd6dc7434268c79195f866718875

Observation da7635fc-2190-4561-b9c3-f2daa71a078a · outbound

This paper cites A Paradigm Shift: The Future of Machine Translation Lies with Large Language Models.

SweEval: Do LLMs Really Swear? A Safety Benchmark for Testing Limits for Enterprise Use A Paradigm Shift: The Future of Machine Translation Lies with Large Language Models

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T14:52:29.088283Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:52:29.088283Z digest=sha256:95187afee95c52921894cb725e0ad9dda0822efd2d3348436047c1324ba484a1

Observation e70cb713-e48b-4f6a-8c0b-d533273190ac · outbound

This paper cites HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust Refusal.

SweEval: Do LLMs Really Swear? A Safety Benchmark for Testing Limits for Enterprise Use HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust Refusal

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T14:52:29.187055Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:52:29.187055Z digest=sha256:ec721ef37e66ccb8df6c4bba9f63e89398bbfd131c18d8cdaa6d7176b6c52567

Observation 96c1bc5c-be2d-424d-b62f-1c8a15e68482 · outbound

This paper cites an unresolved cited work.

SweEval: Do LLMs Really Swear? A Safety Benchmark for Testing Limits for Enterprise Use Unresolved cited work

Reference 29

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:52:32.634095Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T14:52:29.298139Z digest=sha256:28d75b1a1b1ddbda8038bcbf85e7393c9f85ae615ca713b232156712d630d06f

Observation daf4a831-a6f4-4279-b27f-ce3504892117 · outbound

This paper cites an unresolved cited work.

SweEval: Do LLMs Really Swear? A Safety Benchmark for Testing Limits for Enterprise Use Unresolved cited work

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T14:52:29.369006Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:52:29.369006Z digest=sha256:283659a34e833b96ecd92f6abcf2a18ff00eff71612d43ce5154fd482d24fa89

Observation af3517cd-6827-4fff-a2d1-0ce3ba1f690d · outbound

This paper cites an unresolved cited work.

SweEval: Do LLMs Really Swear? A Safety Benchmark for Testing Limits for Enterprise Use Unresolved cited work

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T14:52:29.442168Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:52:29.442168Z digest=sha256:dec9a099be2141e78e46e4c25d48bf27b6f76e33be835a5cd945921a36785624

Observation b52a48f3-7d90-42a4-a53f-80ae4217600b · outbound

This paper cites LLM for Barcodes: Generating Diverse Synthetic Data for Identity Documents.

SweEval: Do LLMs Really Swear? A Safety Benchmark for Testing Limits for Enterprise Use LLM for Barcodes: Generating Diverse Synthetic Data for Identity Documents

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T14:52:29.517727Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:52:29.517727Z digest=sha256:10db1b456455e13dae08e7f780a363ec1354f8cce51ec22697af589ebc6fb33d

Observation fa05beb8-01b6-4ffc-96c4-a948895bb029 · outbound

This paper cites Review of reference generation methods in large language models.

SweEval: Do LLMs Really Swear? A Safety Benchmark for Testing Limits for Enterprise Use Review of reference generation methods in large language models

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T14:52:29.573151Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:52:29.573151Z digest=sha256:d0c4c8e572e5a84aa6a209cbf010ac3aaa1e9d7711b2e5568f42833bab6d1759

Observation 7090315f-e05e-42f2-8158-447c8d1cf6fd · outbound

This paper cites an unresolved cited work.

SweEval: Do LLMs Really Swear? A Safety Benchmark for Testing Limits for Enterprise Use Unresolved cited work

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T14:52:29.624489Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:52:29.624489Z digest=sha256:407736c0870c4d5e979b74abf133d9ea321200adfb5fa31005fc462b8c3baa4e

Observation 5b2fa618-4552-4e31-8336-e1f9e5468b8e · outbound

This paper cites Tokenization Matters: Improving Zero-Shot NER for Indic Languages.

SweEval: Do LLMs Really Swear? A Safety Benchmark for Testing Limits for Enterprise Use Tokenization Matters: Improving Zero-Shot NER for Indic Languages

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T14:52:29.703994Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:52:29.703994Z digest=sha256:2f0b76841c01d5edee9612b625db8d75f042658c3dceeee0f51f4726b5336bed

Observation 686ffdf8-d601-414b-ac05-0c6020e20bd6 · outbound

This paper cites Clinical QA 2.0: Multi-Task Learning for Answer Extraction and Categorization.

SweEval: Do LLMs Really Swear? A Safety Benchmark for Testing Limits for Enterprise Use Clinical QA 2.0: Multi-Task Learning for Answer Extraction and Categorization

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T14:52:29.883990Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:52:29.883990Z digest=sha256:193ff5cbbf6ec4604f6a01289eba3732c9f2c5fc23ca974790f2d995577a2ded

Observation c9cdb51a-fff7-4bf0-8352-94db76df3dcc · outbound

This paper cites Survey of Large Multimodal Model Datasets, Application Categories and Taxonomy.

SweEval: Do LLMs Really Swear? A Safety Benchmark for Testing Limits for Enterprise Use Survey of Large Multimodal Model Datasets, Application Categories and Taxonomy

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T14:52:29.947591Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:52:29.947591Z digest=sha256:3e608a145db4d17cf9c99424facfb99642849642165517ab1dd8c054f3676966

Observation 69935a2f-d505-4a50-b4ac-003d4e18c356 · outbound

This paper cites an unresolved cited work.

SweEval: Do LLMs Really Swear? A Safety Benchmark for Testing Limits for Enterprise Use Unresolved cited work

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T14:52:30.022676Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:52:30.022676Z digest=sha256:3a5efa26c6f121d617b99ea5a66b9acd5f963e44f275a7cb30bf29d4054a40fb

Observation 52b6447c-c493-4b19-89bc-6796518e7df5 · outbound

This paper cites XSTest: A Test Suite for Identifying Exaggerated Safety Behaviours in Large Language Models.

SweEval: Do LLMs Really Swear? A Safety Benchmark for Testing Limits for Enterprise Use XSTest: A Test Suite for Identifying Exaggerated Safety Behaviours in Large Language Models

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T14:52:30.095909Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:52:30.095909Z digest=sha256:ffa01218772b61aa81453b13d5cfe77c61c40a7ff444ac674f5c9bcf06699201

Observation 5058c140-2e80-4716-b643-cb7382fe309d · outbound

This paper cites an unresolved cited work.

SweEval: Do LLMs Really Swear? A Safety Benchmark for Testing Limits for Enterprise Use Unresolved cited work

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T14:52:30.159695Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:52:30.159695Z digest=sha256:c9d414dc6845791e50cd30d69cbab4ba2495193d5dce5f3e4403aa6bb74c3532

Observation 2b1e54c6-b92e-48a4-aefe-4140254a9bb5 · outbound

This paper cites The Language Barrier: Dissecting Safety Challenges of LLMs in Multilingual Contexts.

SweEval: Do LLMs Really Swear? A Safety Benchmark for Testing Limits for Enterprise Use The Language Barrier: Dissecting Safety Challenges of LLMs in Multilingual Contexts

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T14:52:30.239795Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:52:30.239795Z digest=sha256:98dac9568321d4455283afc263faf9ce1650435c293cf4c97b3b2df702a00b5c

Observation 4a3b8bd0-cde7-4ee9-a71d-07972626824b · outbound

This paper cites an unresolved cited work.

SweEval: Do LLMs Really Swear? A Safety Benchmark for Testing Limits for Enterprise Use Unresolved cited work

Reference 43

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:52:32.452987Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T14:52:30.305517Z digest=sha256:754f6332adb5d182217fbc2dd84de8edeedaaa2f033d8c1c91fe8786784acc09

Observation 6875d0ee-2f5a-41e4-a4c4-65e207abcc1f · outbound

This paper cites Text Classification via Large Language Models.

SweEval: Do LLMs Really Swear? A Safety Benchmark for Testing Limits for Enterprise Use Text Classification via Large Language Models

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T14:52:30.375712Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:52:30.375712Z digest=sha256:349135bfbc6ff8ce5af344736608c5cefde2ea4d5ae897c7c55b28ef7b935d3c

Observation 7e4a57dd-0d8c-41dd-b295-f35e2e2bdbe4 · outbound

This paper cites an unresolved cited work.

SweEval: Do LLMs Really Swear? A Safety Benchmark for Testing Limits for Enterprise Use Unresolved cited work

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T14:52:30.469185Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:52:30.469185Z digest=sha256:3de59c07609a0f864c7a97db1456caf1b1f38b70498a0da47c834f4051429b41

Observation 9ad2ab34-9314-4d9f-966e-fba96dd00264 · outbound

This paper cites ALERT: A Comprehensive Benchmark for Assessing Large Language Models' Safety through Red Teaming.

SweEval: Do LLMs Really Swear? A Safety Benchmark for Testing Limits for Enterprise Use ALERT: A Comprehensive Benchmark for Assessing Large Language Models' Safety through Red Teaming

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T14:52:30.530713Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:52:30.530713Z digest=sha256:0b5c8631ae9e624d0b69e75859c060027e377a2153cebeb9140647d4fa5e3583

Observation fff2e369-c818-4d8d-b4a1-3064e957e2bb · outbound

This paper cites an unresolved cited work.

SweEval: Do LLMs Really Swear? A Safety Benchmark for Testing Limits for Enterprise Use Unresolved cited work

Reference 47

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:52:32.284403Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T14:52:30.572397Z digest=sha256:d4e3669f2e480d4e663c3ecce75b5dbe1175cf7d061e26f9abcec419ea234296

Observation 8f548a86-a549-49d8-bf54-2da948643e91 · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

SweEval: Do LLMs Really Swear? A Safety Benchmark for Testing Limits for Enterprise Use LLaMA: Open and Efficient Foundation Language Models

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T14:52:30.616707Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:52:30.616707Z digest=sha256:c153d9a7560129414debdd0b829be97b6fcfd923f5925b42523a427c74c5a5a0

Observation 5661c39f-23e3-44c4-8c81-356c52133435 · outbound

This paper cites All Languages Matter: On the Multilingual Safety of Large Language Models.

SweEval: Do LLMs Really Swear? A Safety Benchmark for Testing Limits for Enterprise Use All Languages Matter: On the Multilingual Safety of Large Language Models

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T14:52:30.662928Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:52:30.662928Z digest=sha256:5d82bff5fa9c554b6ded01efb797b4cf24628e736a76ce2befaf3b66a5b655ee

Observation 8dea832e-a657-4a75-9287-1323bcb16dd5 · outbound

This paper cites Do-Not-Answer: A Dataset for Evaluating Safeguards in LLMs.

SweEval: Do LLMs Really Swear? A Safety Benchmark for Testing Limits for Enterprise Use Do-Not-Answer: A Dataset for Evaluating Safeguards in LLMs

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T14:52:30.727155Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:52:30.727155Z digest=sha256:dc2cb13e2ad7df66001b44513697f08cc28664a57b6a7299e0f4516b786924f9

Observation a4d6c8c2-935d-472d-a3e9-0023f6223a01 · outbound

This paper cites Adaptable and Reliable Text Classification using Large Language Models.

SweEval: Do LLMs Really Swear? A Safety Benchmark for Testing Limits for Enterprise Use Adaptable and Reliable Text Classification using Large Language Models

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-07T14:52:30.798154Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:52:30.798154Z digest=sha256:9f1bf2379a8c6be5839f634a5b4d4f53d0625347f0474112a5ce4b350be393b0

Observation c99df6db-c545-4525-8675-69f32fd72396 · outbound

This paper cites an unresolved cited work.

SweEval: Do LLMs Really Swear? A Safety Benchmark for Testing Limits for Enterprise Use Unresolved cited work

Reference 52

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:52:32.143629Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T14:52:30.869862Z digest=sha256:f2feb872146114fd299a70b4ae5c233e6aa8698513218c97d0b76661cf6eba53

Observation fa468d64-aa82-41e4-93b2-5c3567b007d3 · outbound

This paper cites an unresolved cited work.

SweEval: Do LLMs Really Swear? A Safety Benchmark for Testing Limits for Enterprise Use Unresolved cited work

Reference 53

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:52:31.979578Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T14:52:30.926313Z digest=sha256:07e4ff834b804538240cec7b58bf1e823364e1f8f8e5579ab8d5f558694ac49d

Observation 6d811496-4a3b-4978-864f-92c658110c44 · outbound

This paper cites SORRY-Bench: Systematically Evaluating Large Language Model Safety Refusal.

SweEval: Do LLMs Really Swear? A Safety Benchmark for Testing Limits for Enterprise Use SORRY-Bench: Systematically Evaluating Large Language Model Safety Refusal

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-07T14:52:30.965756Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:52:30.965756Z digest=sha256:5c28396ce1db3c2c0ef34852daf87c50d28a73d09a0b2dc739bcaf68b689f058

Observation 4b32d5de-e91b-44ca-bec3-3bf79eb58ad6 · outbound

This paper cites an unresolved cited work.

SweEval: Do LLMs Really Swear? A Safety Benchmark for Testing Limits for Enterprise Use Unresolved cited work

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-07T14:52:31.008092Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:52:31.008092Z digest=sha256:02bd08287c52d250c973b5a6d40a601f14cb01f44ec2fe3f2ea12454742b2a8e

Observation 6df9c3e1-659f-4b29-9f90-bbf5a8fe15e0 · outbound

This paper cites an unresolved cited work.

SweEval: Do LLMs Really Swear? A Safety Benchmark for Testing Limits for Enterprise Use Unresolved cited work

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-07T14:52:31.085302Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:52:31.085302Z digest=sha256:9733bcc099efe968689988c407c4ca6c91d19babda8182fac494d79fe96fcb9a

Observation a444735e-47da-4991-9cee-16ece478ec49 · outbound

This paper cites SafetyBench: Evaluating the Safety of Large Language Models.

SweEval: Do LLMs Really Swear? A Safety Benchmark for Testing Limits for Enterprise Use SafetyBench: Evaluating the Safety of Large Language Models

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-07T14:52:31.148227Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:52:31.148227Z digest=sha256:347d0a068073efe5bbbab1911e7f400fac910913f0c9a5a24634d0d109dc3412

Observation a1c9744b-1e2d-4831-be5f-977c415f43c2 · outbound

This paper cites an unresolved cited work.

SweEval: Do LLMs Really Swear? A Safety Benchmark for Testing Limits for Enterprise Use Unresolved cited work

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-07T14:52:31.213008Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:52:31.213008Z digest=sha256:406f18d23a761b783ca841ad1160e10fdb165a37276699f5b59e46d7bf11bab3

Observation 419e77c1-3e29-460a-b529-1c910287c98e · outbound

This paper cites Universal and Transferable Adversarial Attacks on Aligned Language Models.

SweEval: Do LLMs Really Swear? A Safety Benchmark for Testing Limits for Enterprise Use Universal and Transferable Adversarial Attacks on Aligned Language Models

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-07T14:52:31.262024Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:52:31.262024Z digest=sha256:b0a75638e336f1b4f1952e58f6abc39e024ada2aae0364361753be8da7f155b4

Pith citing papers

No inbound Pith citation observations are available.