Pith. sign in

Paper Citation Record · LEDGER

SafeSci: Safety Evaluation of Large Language Models in Science Domains and Beyond

As of 2 August 2026, this Paper Citation Record lists 27 of 27 outbound references and 2 inbound Pith citation observations for arXiv:2603.01589.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2603.01589 v2

Coverage vector

measured 27 of 27 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-07-11T11:50:26.030339Z

measured 29 of 29 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-01T06:32:01.292127+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-01T16:21:25.563051Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

27 of 27 outbound references displayed

  • verified exact16
  • verified fuzzy2
  • unresolved7
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 69bccb58-f08e-485b-ba64-87e140781be5 · outbound

This paper cites an unresolved cited work.

SafeSci: Safety Evaluation of Large Language Models in Science Domains and Beyond Unresolved cited work

Reference 1

Resolution
unresolved
raw_fallback, observed 2026-05-15T18:01:26.341189Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-15T18:00:10.534346Z digest=sha256:82497ae1e8154cb2ab8a6b6e1e0ff240a40d88ee0281d38f98eff4f8f26324df

Observation a0463868-eace-4862-a74e-b46ceda5927a · outbound

This paper cites We split all 125 tasks into two sets,Gated Public Access setand HighRisk Restricted Access set.

SafeSci: Safety Evaluation of Large Language Models in Science Domains and Beyond We split all 125 tasks into two sets,Gated Public Access setand HighRisk Restricted Access set

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T18:01:26.345078Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-15T18:00:10.534346Z digest=sha256:139581b20f79358287aaad216ada45b75f853b31bccc0291ac3446f48f722c8e

Observation 50943350-583d-436c-9543-29f2b5d3c95e · outbound

This paper cites an unresolved cited work.

SafeSci: Safety Evaluation of Large Language Models in Science Domains and Beyond Unresolved cited work

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-15T18:01:25.541805Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-15T18:00:10.534346Z digest=sha256:50cf85787dfa239a72251f08821e5062750628a6dabfaebecb7e5f63f8f76edb

Observation b2589ea2-7fa2-4c2d-92ff-72d21cafe9ad · outbound

This paper cites Intern-s1: A scientific multimodal foundation model.

SafeSci: Safety Evaluation of Large Language Models in Science Domains and Beyond Intern-s1: A scientific multimodal foundation model

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-15T18:01:25.535498Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-15T18:00:10.534346Z digest=sha256:642b06492cd23c4e68bb73b6446826c486a40079a42039191823ed91ba7fb96f

Observation 472464fd-c3a2-4519-919b-01ee2029bf4f · outbound

This paper cites Why Should Adversarial Perturbations be Imperceptible? Rethink the Research Paradigm in Adversarial NLP.

SafeSci: Safety Evaluation of Large Language Models in Science Domains and Beyond Why Should Adversarial Perturbations be Imperceptible? Rethink the Research Paradigm in Adversarial NLP

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-15T18:01:25.528515Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-15T18:00:10.534346Z digest=sha256:3fab36832ddb0a26b433b25eb8cd9d1a3827809fb70e44edb2b103df2820782a

Observation 0fe1c598-9fcc-4c8e-916e-a273ed8c76d0 · outbound

This paper cites SuperGPQA: Scaling LLM Evaluation across 285 Graduate Disciplines.

SafeSci: Safety Evaluation of Large Language Models in Science Domains and Beyond SuperGPQA: Scaling LLM Evaluation across 285 Graduate Disciplines

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-16T00:47:04.366365Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-15T18:00:10.534346Z digest=sha256:78b9834e543804c29cd38160e2274fd7d405d4c77632d321514e56e809c73833

Observation 7a1ead7c-9d98-424e-8806-31927ff8f8a4 · outbound

This paper cites The Llama 3 Herd of Models.

SafeSci: Safety Evaluation of Large Language Models in Science Domains and Beyond The Llama 3 Herd of Models

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-05-15T18:01:25.554485Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-15T18:00:10.534346Z digest=sha256:61819d94b46d682d627cc4e4b6bf241afae1d48677286ebb7fb3a946beea5757

Observation a12b6004-4cd2-4988-88ac-fc6ef244fb63 · outbound

This paper cites 2024 , url =.

SafeSci: Safety Evaluation of Large Language Models in Science Domains and Beyond 2024 , url =

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-15T18:01:25.521960Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-15T18:00:10.534346Z digest=sha256:4001aded8a3d1e65c7e250df7fe9c396825070c0a90507a8bc1792b8fd0fbe72

Observation 36667f83-b39d-4a1a-bc61-5cc16d782a26 · outbound

This paper cites an unresolved cited work.

SafeSci: Safety Evaluation of Large Language Models in Science Domains and Beyond Unresolved cited work

Reference 9

Resolution
unresolved
raw_fallback, observed 2026-05-15T18:01:26.336964Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-15T18:00:10.534346Z digest=sha256:9e0469776cef5d01fcff1267e0490cd373aca52eee0fe620059098abc2c0cb0e

Observation 678099fb-4732-426b-b9a1-8f0e79af9e86 · outbound

This paper cites Control Risk for Potential Misuse of Artificial Intelligence in Science.

SafeSci: Safety Evaluation of Large Language Models in Science Domains and Beyond Control Risk for Potential Misuse of Artificial Intelligence in Science

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-15T18:01:25.547934Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-15T18:00:10.534346Z digest=sha256:644dfc20cd31a90c1426f8b5a9032cce0cd979edc423ca18144afdd46aee293f

Observation 905da372-f9b8-4398-9af3-2bda31f53ec4 · outbound

This paper cites an unresolved cited work.

SafeSci: Safety Evaluation of Large Language Models in Science Domains and Beyond Unresolved cited work

Reference 11

Resolution
unresolved
raw_fallback, observed 2026-05-15T18:01:26.332750Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-15T18:00:10.534346Z digest=sha256:8eab3aa2725960d53577ea282026d1fa591a7d4af933cd849ddfef7d2ed5b618

Observation 6716e6ff-3368-4e39-8e49-7a5bc2dbdcbe · outbound

This paper cites SoSBench: Benchmarking Safety Alignment on Six Scientific Domains.

SafeSci: Safety Evaluation of Large Language Models in Science Domains and Beyond SoSBench: Benchmarking Safety Alignment on Six Scientific Domains

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-05-15T18:01:25.616943Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-15T18:00:10.534346Z digest=sha256:6d716ae4b8f0020760c53f56be8fc6d002e9f2a86fbe11dde0a8035689910d6c

Observation 3d135be1-68bf-4279-b346-7668da1352cb · outbound

This paper cites an unresolved cited work.

SafeSci: Safety Evaluation of Large Language Models in Science Domains and Beyond Unresolved cited work

Reference 13

Resolution
unresolved
raw_fallback, observed 2026-05-15T18:01:26.319892Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-15T18:00:10.534346Z digest=sha256:27dd031a1120e23331306f7887e4b3cb065606c227163eac49d5aaa75cba709a

Observation 07eafd48-7b4a-4fcf-8c25-f10411897d8b · outbound

This paper cites an unresolved cited work.

SafeSci: Safety Evaluation of Large Language Models in Science Domains and Beyond Unresolved cited work

Reference 14

Resolution
unresolved
raw_fallback, observed 2026-05-15T18:01:26.315522Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-15T18:00:10.534346Z digest=sha256:9f57543e1bc341e479056d6cdbc29a7e2585cc2f4f2281216faf30442c791b81

Observation 4abfd7fe-5cb2-482c-9ca5-042db629d440 · outbound

This paper cites SafeWork-R1: Coevolving Safety and Intelligence under the AI-45$^{\circ}$ Law.

SafeSci: Safety Evaluation of Large Language Models in Science Domains and Beyond SafeWork-R1: Coevolving Safety and Intelligence under the AI-45$^{\circ}$ Law

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-15T18:01:25.610609Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-15T18:00:10.534346Z digest=sha256:40fee6f2edec187be76f1f1d2d028b941e1b389069a5f866be917cfc6e1b6d9d

Observation c8b38e62-a389-450e-970f-bb5c12691c84 · outbound

This paper cites The WMDP Benchmark: Measuring and Reducing Malicious Use With Unlearning.

SafeSci: Safety Evaluation of Large Language Models in Science Domains and Beyond The WMDP Benchmark: Measuring and Reducing Malicious Use With Unlearning

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-15T19:55:50.014661Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-15T18:00:10.534346Z digest=sha256:7e0eaf3e3b4e8d00e2b9ec65934b953ad85a65c24f5ef858c8895ceb277a1c3c

Observation 7f150f46-66c7-400c-8f0d-d602b34f1a6e · outbound

This paper cites HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust Refusal.

SafeSci: Safety Evaluation of Large Language Models in Science Domains and Beyond HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust Refusal

Reference 18

Resolution
verified exact
local_arxiv, observed 2026-05-15T18:01:25.623098Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-15T18:00:10.534346Z digest=sha256:d837cb70262f6c432f69a0f837882452d87ed2032d5849661948d585838c7af0

Observation 7a9f240b-0e1a-4899-9896-482ff365bc62 · outbound

This paper cites Probing and Steering Evaluation Awareness of Language Models.

SafeSci: Safety Evaluation of Large Language Models in Science Domains and Beyond Probing and Steering Evaluation Awareness of Language Models

Reference 19

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T18:01:25.603359Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-15T18:00:10.534346Z digest=sha256:6a9957f2e9b42b27db327d3abb5ff011bae07df7d34069b060be12da33b58d34

Observation b6022c6a-5560-4088-9723-2352897489c3 · outbound

This paper cites an unresolved cited work.

SafeSci: Safety Evaluation of Large Language Models in Science Domains and Beyond Unresolved cited work

Reference 20

Resolution
unresolved
raw_fallback, observed 2026-05-15T18:01:26.328132Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-15T18:00:10.534346Z digest=sha256:78f71881ee90370ffa05b11a2a5a6f719aec9fe0bb0dff78c6c5bc2e57aa05b7

Observation 614640aa-d0a4-43d6-bf38-ab4e4f7d7006 · outbound

This paper cites an unresolved cited work.

SafeSci: Safety Evaluation of Large Language Models in Science Domains and Beyond Unresolved cited work

Reference 21

Resolution
unresolved
raw_fallback, observed 2026-05-15T18:01:26.323679Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-15T18:00:10.534346Z digest=sha256:44809767a9da934db38b4f688c58f2b149cb72930961a754fde205e16b99babb

Observation 6bb84f98-db4b-472d-beeb-68fab9f77420 · outbound

This paper cites Qwen3 Technical Report.

SafeSci: Safety Evaluation of Large Language Models in Science Domains and Beyond Qwen3 Technical Report

Reference 22

Resolution
verified exact
local_arxiv, observed 2026-05-15T18:01:25.596110Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-15T18:00:10.534346Z digest=sha256:1627ea746bbbfdc0f47b9e9782882eda5c2926ea0381d5f087dee65bef8ba099

Observation cee15327-83a4-414a-b92b-f201f669c678 · outbound

This paper cites Accessed: 2026-01-29.

SafeSci: Safety Evaluation of Large Language Models in Science Domains and Beyond Accessed: 2026-01-29

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T18:01:26.310950Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-15T18:00:10.534346Z digest=sha256:3c40f49f21d262ca15e4b598405c8062e9e23e06b2e4cd255ca0facaed4a874e

Observation 93886e3c-a0e1-4be5-8bce-e27eeb57d4ac · outbound

This paper cites Qwen3 Embedding: Advancing Text Embedding and Reranking Through Foundation Models.

SafeSci: Safety Evaluation of Large Language Models in Science Domains and Beyond Qwen3 Embedding: Advancing Text Embedding and Reranking Through Foundation Models

Reference 24

Resolution
verified exact
local_arxiv, observed 2026-05-15T18:01:25.243853Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-07-11T11:50:26.030339Z digest=sha256:ccc744d0efe22909516cbe98318c58026ac05b6d4a1cf246d960cd8faaba6755

Observation dbee5bae-42ba-40d1-a12f-3e827a123fd9 · outbound

This paper cites SafetyFlow: An Agent-Flow System for Automated LLM Safety Benchmarking.

SafeSci: Safety Evaluation of Large Language Models in Science Domains and Beyond SafetyFlow: An Agent-Flow System for Automated LLM Safety Benchmarking

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-15T18:01:25.575015Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-15T18:00:10.534346Z digest=sha256:c0d26cfc8efe23357f69663ceb2100c24158e9e9223f3022f13377d2e31ec445

Observation f6f258be-97c4-4423-ad54-4beea4f501d0 · outbound

This paper cites Universal and Transferable Adversarial Attacks on Aligned Language Models.

SafeSci: Safety Evaluation of Large Language Models in Science Domains and Beyond Universal and Transferable Adversarial Attacks on Aligned Language Models

Reference 26

Resolution
verified exact
local_arxiv, observed 2026-05-15T18:01:25.588604Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-15T18:00:10.534346Z digest=sha256:8a8db7917dd4dfacba4187da81094e3e44c07c29c4bbf71c2f9153985fe6c719

Observation af3b1add-78f6-45e4-a25c-52dd2c13dea4 · outbound

This paper cites Ans.” represent if the questions have corresponding answers. “Rep.

SafeSci: Safety Evaluation of Large Language Models in Science Domains and Beyond Ans.” represent if the questions have corresponding answers. “Rep

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-15T18:01:25.581771Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-15T18:00:10.534346Z digest=sha256:84f93498e974ea92c258c87d69867c32eda60403004be78756938cd56def6d03

Observation 0fe6894c-4ced-4bb3-82ff-d16ee20b6920 · outbound

This paper cites an unresolved cited work.

SafeSci: Safety Evaluation of Large Language Models in Science Domains and Beyond Unresolved cited work

Reference 28

Resolution
malformed identifier
raw_fallback, observed 2026-05-15T18:01:26.349334Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-15T18:00:10.534346Z digest=sha256:3c807a8c3684f9ee9265157b53ccc04634164013b13bc0651c780ceb8cf992e6

Pith citing papers

Observation affd77b5-38e3-41fb-8a5b-5339ac7598ec · inbound

An Early Warning of Emerging Biosecurity Risks in Frontier LLMs cites this paper.

An Early Warning of Emerging Biosecurity Risks in Frontier LLMs SafeSci: Safety Evaluation of Large Language Models in Science Domains and Beyond

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-01T16:21:25.563051Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T16:21:25.563051Z digest=sha256:50516026e57ee75d2d57a9b3bb83afb13386b42fada967ea087ae8e698c17430

Observation dfb99e88-fc71-4050-b015-9d7219c9c2a6 · inbound

SciHazard: A Benchmark for Measuring Scientific Safety Risks with Decomposed Harm Scoring cites this paper.

SciHazard: A Benchmark for Measuring Scientific Safety Risks with Decomposed Harm Scoring SafeSci: Safety Evaluation of Large Language Models in Science Domains and Beyond

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-01T14:45:44.097098Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T14:45:44.097098Z digest=sha256:3fa01f4c228940f8f23b1f562d05404eb9ef03878daad43197061368c5e2eb96