Pith. sign in

Paper Citation Record · LEDGER

S-Eval: Towards Automated and Comprehensive Safety Evaluation for Large Language Models

As of 19 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 16 inbound Pith citation observations for arXiv:2405.14191.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2405.14191 v4

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 16 of 16 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 16 of 16 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T10:12:00.899754Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T18:10:00.725685Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 246381f3-539f-4c3e-a52a-21eafb6c6b56 · inbound

From Local to Global: A Graph RAG Approach to Query-Focused Summarization cites this paper.

From Local to Global: A Graph RAG Approach to Query-Focused Summarization S-Eval: Towards Automated and Comprehensive Safety Evaluation for Large Language Models

Reference 74

Resolution
verified exact
arxiv_id, observed 2026-08-04T01:57:26.797777Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-05-11T05:10:57.816312Z digest=sha256:58d4daa73fd75496a0b567f0a4fca3519f6e547f329026790a0efd84378400b9

Observation 627f8aad-0086-494d-b7bf-38053d7e8e8c · inbound

From Allies to Adversaries: Manipulating LLM Tool-Calling through Adversarial Injection cites this paper.

From Allies to Adversaries: Manipulating LLM Tool-Calling through Adversarial Injection S-Eval: Towards Automated and Comprehensive Safety Evaluation for Large Language Models

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-11T16:17:43.479613Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:17:43.479613Z digest=sha256:1c29faa1ded8192a5a95c7f957e192df37f9d9155feb4dbaf555ccae71c16a96

Observation 3c2d03b5-39a7-4daf-9977-694f7bab023b · inbound

A Method for Enhancing the Safety of Large Model Generation Based on Multi-dimensional Attack and Defense cites this paper.

A Method for Enhancing the Safety of Large Model Generation Based on Multi-dimensional Attack and Defense S-Eval: Towards Automated and Comprehensive Safety Evaluation for Large Language Models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-10T22:53:20.842188Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:53:20.842188Z digest=sha256:848182510e1dc306241c655702e08f5f10a32f7e50197ded528275f472d71f32

Observation b794624a-4476-4186-9a07-f96a5c272ab1 · inbound

Value Compass Benchmarks: A Platform for Fundamental and Validated Evaluation of LLMs Values cites this paper.

Value Compass Benchmarks: A Platform for Fundamental and Validated Evaluation of LLMs Values S-Eval: Towards Automated and Comprehensive Safety Evaluation for Large Language Models

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-10T20:53:59.898841Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:53:59.898841Z digest=sha256:2e6a562d41399ccdc1a82b4dfda246323aa48bb39a0f54e1a3d82edc80398d3e

Observation 2b96bf4d-ed93-4bed-b67b-095f5ad09327 · inbound

Early External Safety Testing of OpenAI's o3-mini: Insights from the Pre-Deployment Evaluation cites this paper.

Early External Safety Testing of OpenAI's o3-mini: Insights from the Pre-Deployment Evaluation S-Eval: Towards Automated and Comprehensive Safety Evaluation for Large Language Models

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-10T04:37:16.948537Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T04:37:16.948537Z digest=sha256:c9248e7785c46e2b77553fdc6198759989aca982e16920ef7f9986898c1a72b1

Observation ed2b1a91-fe68-4fd9-8dba-084e3064c9d9 · inbound

o3-mini vs DeepSeek-R1: Which One is Safer? cites this paper.

o3-mini vs DeepSeek-R1: Which One is Safer? S-Eval: Towards Automated and Comprehensive Safety Evaluation for Large Language Models

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-09T23:32:16.751706Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T23:32:16.751706Z digest=sha256:c6d541e6929f51550e7c5350361a2a7cd631a26519fe95e1fa1c5013a0041e29

Observation 87a1d1c7-2a9d-4572-9231-8c09383047a7 · inbound

Psychometric-Based Evaluation for Theorem Proving with Large Language Models cites this paper.

Psychometric-Based Evaluation for Theorem Proving with Large Language Models S-Eval: Towards Automated and Comprehensive Safety Evaluation for Large Language Models

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-09T17:36:09.653389Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T17:36:09.653389Z digest=sha256:036dc46c208df70501804ffa0e936536e16edbbb59bafe5da2883a82a07eeb64

Observation 545f1f60-e197-4e43-bff9-c3d1015fdb0b · inbound

The Science of Evaluating Foundation Models cites this paper.

The Science of Evaluating Foundation Models S-Eval: Towards Automated and Comprehensive Safety Evaluation for Large Language Models

Reference 104

Resolution
unresolved
no resolver link, observed 2026-08-07T23:35:42.916246Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:35:42.916246Z digest=sha256:599194fb1f9a2d245ca9dc2e4ee423581112b295e9ee7380e007ddfbd5e8271c

Observation aa10d285-7172-43b9-88a1-d48d23867685 · inbound

Toward Generalizable Evaluation in the LLM Era: A Survey Beyond Benchmarks cites this paper.

Toward Generalizable Evaluation in the LLM Era: A Survey Beyond Benchmarks S-Eval: Towards Automated and Comprehensive Safety Evaluation for Large Language Models

Reference 179

Resolution
unresolved
no resolver link, observed 2026-08-16T10:12:00.899754Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:12:00.899754Z digest=sha256:312d6e80cacab1c75442f9028280a42f1a19f990ba90fc4a34ec94c84dd66887

Observation 2fe61643-fdf3-4e05-9533-fff34d8e8f5d · inbound

The Aloe Family Recipe for Open and Specialized Healthcare LLMs cites this paper.

The Aloe Family Recipe for Open and Specialized Healthcare LLMs S-Eval: Towards Automated and Comprehensive Safety Evaluation for Large Language Models

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-15T23:36:30.053396Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:36:30.053396Z digest=sha256:19b0afba46cfa3d6a1b3ce06bc1808a8006265cd306dc06d7bf980502b7f8278

Observation d4ac6a66-df10-4627-9818-94933ba03a0e · inbound

The Scales of Justitia: A Comprehensive Survey on Safety Evaluation of LLMs cites this paper.

The Scales of Justitia: A Comprehensive Survey on Safety Evaluation of LLMs S-Eval: Towards Automated and Comprehensive Safety Evaluation for Large Language Models

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-07T10:17:26.844499Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:17:26.844499Z digest=sha256:a54c5915a2d3e0a16cc6d22a53e6f3808a9414a1dd174938e572f1ac3c95ad69

Observation 303efba4-6986-4882-ab4a-783b4a3ba1d2 · inbound

ORFuzz: Fuzzing the "Other Side" of LLM Safety -- Testing Over-Refusal cites this paper.

ORFuzz: Fuzzing the "Other Side" of LLM Safety -- Testing Over-Refusal S-Eval: Towards Automated and Comprehensive Safety Evaluation for Large Language Models

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-08-04T01:57:26.797777Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-18T23:27:52.438709Z digest=sha256:b3824e3380952bba786ac14e0ad84af01da502d7b59006fc80664846d9b85bad

Observation d7650cbd-c31f-4368-ad84-0f9410e5f121 · inbound

Controlling the Risk of Corrupted Contexts for Language Models via Early-Exiting cites this paper.

Controlling the Risk of Corrupted Contexts for Language Models via Early-Exiting S-Eval: Towards Automated and Comprehensive Safety Evaluation for Large Language Models

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-04T12:45:26.905184Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T12:45:26.905184Z digest=sha256:20576f5bac7f24d6a10f7fb47106759ff9908643634ae0011113dab319be4b1e

Observation 1a5121e0-d0c8-4202-a219-9821cb75e50e · inbound

Multilingual Refusal Alignment for Safer Large Language Models cites this paper.

Multilingual Refusal Alignment for Safer Large Language Models S-Eval: Towards Automated and Comprehensive Safety Evaluation for Large Language Models

Reference 72

Resolution
verified exact
arxiv_id, observed 2026-08-04T01:57:26.797777Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-07-04T18:09:17.531727Z digest=sha256:79858ea7a36199571ac517038ecefc943a89e2c84e85a57888580b6f58e046fe

Observation d42ba51f-292d-4d59-8f56-08e25b9b7ef9 · inbound

Yuvion LLM: An Adversarially-Aware Large Language Model for Content And AI Safety cites this paper.

Yuvion LLM: An Adversarially-Aware Large Language Model for Content And AI Safety S-Eval: Towards Automated and Comprehensive Safety Evaluation for Large Language Models

Reference 38

Resolution
verified exact
arxiv_id, observed 2026-08-04T01:57:26.797777Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-29T00:46:03.210076Z digest=sha256:b737f09f50488cdf92a725e282fdcc10633e0990cac4ee73a98be732b0bc8278

Observation 077f18ff-e8be-4268-a724-7044051ea787 · inbound

When Refusal Looks Safe: The Refusal-Cue Shortcut in Safety Guard Models cites this paper.

When Refusal Looks Safe: The Refusal-Cue Shortcut in Safety Guard Models S-Eval: Towards Automated and Comprehensive Safety Evaluation for Large Language Models

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-05T23:46:20.542247Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:46:20.542247Z digest=sha256:210da692abd1dd19787ec3705c67d7e18a5b6ffad4a0087054d98d5025052a39