Pith. sign in

Paper Citation Record · LEDGER

ChemSafetyBench: Benchmarking LLM Safety on Chemistry Domain

As of 18 August 2026, this Paper Citation Record lists 51 of 51 outbound references and 8 inbound Pith citation observations for arXiv:2411.16736.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2411.16736 v1

Coverage vector

measured 51 of 51 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T14:13:30.618840Z

measured 59 of 59 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 8 of 8 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T22:31:28.922399Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-25T06:06:43.111814Z

Reference resolution

51 of 51 outbound references displayed

  • verified exact0
  • verified fuzzy20
  • unresolved30
  • parse uncertain1
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 725413de-4f3b-47db-8a32-e6b56e87598d · outbound

This paper cites GPT-4 Technical Report.

ChemSafetyBench: Benchmarking LLM Safety on Chemistry Domain GPT-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-12T14:13:29.573144Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:13:29.573144Z digest=sha256:d94ade3ec9736003d5cd334e283509e68accf1c36be2771d7040e7ccbb5cbe9a

Observation c73b7420-3bbe-481d-b3b7-334844a7032e · outbound

This paper cites an unresolved cited work.

ChemSafetyBench: Benchmarking LLM Safety on Chemistry Domain Unresolved cited work

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-12T14:13:29.597406Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:13:29.597406Z digest=sha256:0ec7c31f6784c00aad97510626fa6474dc340c2f0ae4a97904501600a193244d

Observation e5cc7824-c022-4afc-9028-1391563cd134 · outbound

This paper cites Llama 3 model card.

ChemSafetyBench: Benchmarking LLM Safety on Chemistry Domain Llama 3 model card

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-12T14:13:29.604398Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:13:29.604398Z digest=sha256:a19d044a2c5618c1d61c33c012102dae44b4cb3c7554e93bc26613b23ef4fe28

Observation 56c0175f-2563-4227-82d4-cf62b32ef8c8 · outbound

This paper cites The claude 3 model family: Opus, sonnet, haiku.

ChemSafetyBench: Benchmarking LLM Safety on Chemistry Domain The claude 3 model family: Opus, sonnet, haiku

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-12T14:13:29.610589Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:13:29.610589Z digest=sha256:43894156207db932840d32c0353fbbeba73e4f9be5e5da6f7bf163aaa0c837c6

Observation 286ea39e-a229-4258-861a-65161cdf1422 · outbound

This paper cites Qwen Technical Report.

ChemSafetyBench: Benchmarking LLM Safety on Chemistry Domain Qwen Technical Report

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-12T14:13:29.633261Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:13:29.633261Z digest=sha256:8d38e0b1f23fe0d4b1178d3d6901117cb32245580b1c1a630207ea9f0d10066c

Observation 6799db45-a564-423c-a4a5-04d43bcf6744 · outbound

This paper cites Autonomous chemical research with large language models.

ChemSafetyBench: Benchmarking LLM Safety on Chemistry Domain Autonomous chemical research with large language models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-12T14:13:29.674124Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:13:29.674124Z digest=sha256:5f91d74f6f107ba2748df1ac1e7d7fef5865eb2d436448480c30cea6d0d1fb6a

Observation 4b2ec2bb-5528-4b0f-8545-86056f5c2f64 · outbound

This paper cites Manning, and Roxana Daneshjou.

ChemSafetyBench: Benchmarking LLM Safety on Chemistry Domain Manning, and Roxana Daneshjou

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:13:32.656089Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T14:13:29.691608Z digest=sha256:c4a2f9a6c0a1f1bc65a7fc295f7e984ee6d2db2c119800b694382b6f5249d459

Observation f2321ce5-37fc-4d13-8af0-31e70ba03428 · outbound

This paper cites Bioinfo-bench: A simple benchmark framework for llm bioinformatics skills evaluation.

ChemSafetyBench: Benchmarking LLM Safety on Chemistry Domain Bioinfo-bench: A simple benchmark framework for llm bioinformatics skills evaluation

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:13:32.436165Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T14:13:29.697690Z digest=sha256:d977161f5e29e4c3e4a38e823d935997ac23ccd855c8ba4167e2be3d0fa310a0

Observation bc27f439-5fe0-421e-977d-03269b8582a1 · outbound

This paper cites 7 revealing ways ais fail: Neural networks can be disastrously brittle, forgetful, and surprisingly bad at math.

ChemSafetyBench: Benchmarking LLM Safety on Chemistry Domain 7 revealing ways ais fail: Neural networks can be disastrously brittle, forgetful, and surprisingly bad at math

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:13:32.407480Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T14:13:29.703002Z digest=sha256:ae202845e271c4e1b8c6dbd3beb9bb24904da46a689acb92f27f852097f176eb

Observation 92ce51a7-7600-4755-add1-36fd4f4e2828 · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

ChemSafetyBench: Benchmarking LLM Safety on Chemistry Domain Training Verifiers to Solve Math Word Problems

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-12T14:13:29.709737Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:13:29.709737Z digest=sha256:f90ef84d85b632ce9222a196c9ea5622ff70e772454770ee2d28ba2692622cba

Observation dcc1e499-e7a1-4f17-9e4a-aac80db084e5 · outbound

This paper cites Reaxys, 2024.

ChemSafetyBench: Benchmarking LLM Safety on Chemistry Domain Reaxys, 2024

Reference 11

Resolution
parse uncertain
raw_fallback, observed 2026-08-12T14:13:32.385695Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T14:13:29.716484Z digest=sha256:6ea5ac71cbfc5026d1be51d02718bd258dc8b32421465b8b4278f80e5bc03072

Observation 603757da-0c17-470d-9cf5-5c377758aa20 · outbound

This paper cites Pubchem, 2024.

ChemSafetyBench: Benchmarking LLM Safety on Chemistry Domain Pubchem, 2024

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:13:32.361507Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T14:13:29.722191Z digest=sha256:99c98c342d9572a217b1eede2732451beaecc52d51aab10858cac54f9c3d5970

Observation 862b633f-002e-46d1-a712-e0574b4be724 · outbound

This paper cites What can large language models do in chemistry? a comprehensive benchmark on eight tasks.

ChemSafetyBench: Benchmarking LLM Safety on Chemistry Domain What can large language models do in chemistry? a comprehensive benchmark on eight tasks

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:13:32.337335Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T14:13:29.753956Z digest=sha256:f0e753647f1356808861a2b0fba928678fd8402422a75e566526dced2ecdbb86

Observation 96573458-b7e6-437b-bbcf-409c957b5d8e · outbound

This paper cites Testing llm performance on the physics gre: some observations, 2023.

ChemSafetyBench: Benchmarking LLM Safety on Chemistry Domain Testing llm performance on the physics gre: some observations, 2023

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:13:32.313049Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T14:13:29.801405Z digest=sha256:0d6d91267bf09c27ec9db12f71f6e0096874b2db59711722cf0718694a779b19

Observation 2a72e99d-6cef-482d-b746-0ca9d31d536a · outbound

This paper cites A survey on large language models: Applications, challenges, limitations, and practical usage.

ChemSafetyBench: Benchmarking LLM Safety on Chemistry Domain A survey on large language models: Applications, challenges, limitations, and practical usage

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-12T14:13:29.806063Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:13:29.806063Z digest=sha256:806aeeb920e7a99bc850c0c9fdc6f3bcc058dc4f7dd4dcc8abbd34a7a01b9af5

Observation 9ec6be15-c050-440e-96a4-388f71cc675b · outbound

This paper cites Control risk for potential misuse of artificial intelligence in science, 2023.

ChemSafetyBench: Benchmarking LLM Safety on Chemistry Domain Control risk for potential misuse of artificial intelligence in science, 2023

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:13:32.152244Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T14:13:29.810841Z digest=sha256:1b2d9b064e28c6e1fa9056117924902b6e8a69f89f61c3873ea39173cbb1f588

Observation b0f3e916-d60f-4496-a253-6d652f7d43a6 · outbound

This paper cites Measuring Mathematical Problem Solving With the MATH Dataset.

ChemSafetyBench: Benchmarking LLM Safety on Chemistry Domain Measuring Mathematical Problem Solving With the MATH Dataset

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-12T14:13:29.816346Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:13:29.816346Z digest=sha256:83edc15043d03bbe349604574d24a6dfc094ff342cebbc6487b69fbc199fbf0b

Observation 78ea54ef-4993-4424-9b62-327f83ec94b7 · outbound

This paper cites Amortizing intractable inference in large language models.

ChemSafetyBench: Benchmarking LLM Safety on Chemistry Domain Amortizing intractable inference in large language models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-12T14:13:29.822476Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:13:29.822476Z digest=sha256:681addc60fbe52e03cc117e8ca22dff6cc52acf12d06a8322ba6dd606e5c1879

Observation dbebfee4-bf42-40d0-b12a-e4f429c54186 · outbound

This paper cites Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations.

ChemSafetyBench: Benchmarking LLM Safety on Chemistry Domain Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-12T14:13:29.831623Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:13:29.831623Z digest=sha256:41f9cef8e0a6e759dbd667a6c8f1fdb793699cb84ad622054a341b044550c649

Observation 22ea5419-1a3e-4f19-a983-86722423f3c1 · outbound

This paper cites Leverag- ing large language models for predictive chemistry.

ChemSafetyBench: Benchmarking LLM Safety on Chemistry Domain Leverag- ing large language models for predictive chemistry

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:13:31.988696Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T14:13:29.837221Z digest=sha256:0d9464e541cefe68526e5a5f6c485ad4c79b5e586dc810132d43689c66076cbe

Observation 372d6eb8-33bc-4048-931c-e419cb0d4501 · outbound

This paper cites A compre- hensive evaluation of large language models on benchmark biomedical text processing tasks.

ChemSafetyBench: Benchmarking LLM Safety on Chemistry Domain A compre- hensive evaluation of large language models on benchmark biomedical text processing tasks

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:13:31.950399Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T14:13:29.878291Z digest=sha256:e10a71c90d7634061500aaa46858fe6263eff22248a76195dc9420c7a7e1bd6f

Observation c3a09434-197a-43db-8c09-1321c0593d69 · outbound

This paper cites an unresolved cited work.

ChemSafetyBench: Benchmarking LLM Safety on Chemistry Domain Unresolved cited work

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-12T14:13:29.912941Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:13:29.912941Z digest=sha256:80590f7161f1ce10631442611dd52a12726c85fa1af5120074f38ef2b5f79c44

Observation 68a7c90c-a619-49eb-900f-684a2da1fcf4 · outbound

This paper cites Llm based biological named entity recognition from scientific literature.

ChemSafetyBench: Benchmarking LLM Safety on Chemistry Domain Llm based biological named entity recognition from scientific literature

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-12T14:13:29.919876Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:13:29.919876Z digest=sha256:d6056cd2ab86c63ae99ab43674e41b90b7ad55fdd8eacb8236cf3990d4127577

Observation d0377570-475b-43ca-b87d-cb35a6b8dd0c · outbound

This paper cites Large Language Models in Plant Biology.

ChemSafetyBench: Benchmarking LLM Safety on Chemistry Domain Large Language Models in Plant Biology

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-12T14:13:29.928166Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:13:29.928166Z digest=sha256:297dd6a6bdc3d623c523a0e626d0c8f7a9e1b64419f91f0d9607327b820462a1

Observation 90d8e3c8-4f2f-4112-b717-4f6ae335f782 · outbound

This paper cites SALAD-Bench: A Hierarchical and Comprehensive Safety Benchmark for Large Language Models.

ChemSafetyBench: Benchmarking LLM Safety on Chemistry Domain SALAD-Bench: A Hierarchical and Comprehensive Safety Benchmark for Large Language Models

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-12T14:13:29.934618Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:13:29.934618Z digest=sha256:a18a6d3c8c56b4f10c652aad5ab562b49efec8e8429080d86aea810be8f22fe1

Observation 363eced4-2ff4-43d5-8778-2595877ebb02 · outbound

This paper cites Learn to explain: Multimodal reasoning via thought chains for science question answering.

ChemSafetyBench: Benchmarking LLM Safety on Chemistry Domain Learn to explain: Multimodal reasoning via thought chains for science question answering

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-12T14:13:29.941544Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:13:29.941544Z digest=sha256:61f9f0fbe6c8cd58136efcf4af152022b440814f96659045678833c94284b148

Observation 283a0f19-1621-421c-98c2-33ccecf9eeef · outbound

This paper cites Bran, Sam Cox, Oliver Schilter, Carlo Baldassari, Andrew D White, and Philippe Schwaller.

ChemSafetyBench: Benchmarking LLM Safety on Chemistry Domain Bran, Sam Cox, Oliver Schilter, Carlo Baldassari, Andrew D White, and Philippe Schwaller

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:13:31.876810Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T14:13:29.950036Z digest=sha256:6e43cb6675fcf943d21b1c51a9b5629d257dc0c997d3f1cde606d8671b43b3e6

Observation 1286dc47-0650-45b9-b833-9a5a667c6282 · outbound

This paper cites Feasibility study on parameter adjustment for a humanoid using llm tailoring physical care.

ChemSafetyBench: Benchmarking LLM Safety on Chemistry Domain Feasibility study on parameter adjustment for a humanoid using llm tailoring physical care

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:13:31.815541Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T14:13:29.955973Z digest=sha256:8e129397d563ee743ece58dfc49e4b64011ec4041f7bbc31ccbacc545f268eac

Observation aa4f2c8a-0342-449e-9f8c-483e2a8c807b · outbound

This paper cites Ghs classification results, 2024.

ChemSafetyBench: Benchmarking LLM Safety on Chemistry Domain Ghs classification results, 2024

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:13:31.730801Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T14:13:29.985121Z digest=sha256:14766b7a251591f46051ff5db8f431f9484efb329a6cd370e39c3b0f715464f4

Observation 3b55005b-1c68-48cd-b89f-0506e4a0f26e · outbound

This paper cites Forbidden materials, 2024.

ChemSafetyBench: Benchmarking LLM Safety on Chemistry Domain Forbidden materials, 2024

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:13:31.655625Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T14:13:30.035469Z digest=sha256:7ad811428fddb2eeeb46f639dced4800061b626ad5a0f7884dc68a5ed71a151b

Observation 954e25ab-33c8-4929-b8f1-380b590e7ecf · outbound

This paper cites SafetyPrompts: a Systematic Review of Open Datasets for Evaluating and Improving Large Language Model Safety.

ChemSafetyBench: Benchmarking LLM Safety on Chemistry Domain SafetyPrompts: a Systematic Review of Open Datasets for Evaluating and Improving Large Language Model Safety

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-12T14:13:30.092281Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:13:30.092281Z digest=sha256:d062a87cbe4b2d499cf001a4d0b879a38fc4c44134a132087e7e5e6c2929ab16

Observation dafe91d1-599d-4f26-a3f6-5c41ef31160c · outbound

This paper cites LLM-Prop: Predicting Physical And Electronic Properties Of Crystalline Solids From Their Text Descriptions.

ChemSafetyBench: Benchmarking LLM Safety on Chemistry Domain LLM-Prop: Predicting Physical And Electronic Properties Of Crystalline Solids From Their Text Descriptions

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-12T14:13:30.100549Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:13:30.100549Z digest=sha256:1a8d572f87f5de878f3401dfa75f79fdd1e5015b2c0a5b278d9d4c64411f50f0

Observation 2b384e7b-4249-4d34-917a-c43584a04316 · outbound

This paper cites Sci- enceqa: A novel resource for question answering on scholarly articles.

ChemSafetyBench: Benchmarking LLM Safety on Chemistry Domain Sci- enceqa: A novel resource for question answering on scholarly articles

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-12T14:13:30.105680Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:13:30.105680Z digest=sha256:0623ff929b9aa4655dcd90a87a78da4322fff1db78f47e4333279b95c447c5f3

Observation e8cebf45-4ba9-44e4-9a95-bca374ce03c5 · outbound

This paper cites Scifinder, 2024.

ChemSafetyBench: Benchmarking LLM Safety on Chemistry Domain Scifinder, 2024

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:13:31.476645Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T14:13:30.111553Z digest=sha256:45826d8c40df95decb77b48aee5d00544ed7337d2a94ad95959af451e5c7218d

Observation c4d34305-8f25-421a-a737-47ba038d9d36 · outbound

This paper cites Large language models encode clinical knowledge.

ChemSafetyBench: Benchmarking LLM Safety on Chemistry Domain Large language models encode clinical knowledge

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-12T14:13:30.116533Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:13:30.116533Z digest=sha256:7315d1c36a967b0875cf2c9a2606e1772ff17a90e4dd0e1f23d903a12ee9214c

Observation 3256f40b-4469-4936-8317-a2ed30ce3d90 · outbound

This paper cites Scieval: A multi-level large language model evaluation benchmark for scientific research, 2023.

ChemSafetyBench: Benchmarking LLM Safety on Chemistry Domain Scieval: A multi-level large language model evaluation benchmark for scientific research, 2023

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:13:31.390532Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T14:13:30.122858Z digest=sha256:01c2af86a58cdb333203af0b7768e02b91937eb9e228ee8500838432a85258f5

Observation 29247775-4c1e-4820-bf7d-f694f3e14dfe · outbound

This paper cites Prioritizing safeguarding over autonomy: Risks of llm agents for science, 2024.

ChemSafetyBench: Benchmarking LLM Safety on Chemistry Domain Prioritizing safeguarding over autonomy: Risks of llm agents for science, 2024

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:13:31.371688Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T14:13:30.225702Z digest=sha256:7036ef8bf9c6713bdd8e36f4274f39e6cddb47b67b28d2323a4a4f216d11e343

Observation 0acf68a5-75b1-4a76-8a61-9211234ece49 · outbound

This paper cites ALERT: A Comprehensive Benchmark for Assessing Large Language Models' Safety through Red Teaming.

ChemSafetyBench: Benchmarking LLM Safety on Chemistry Domain ALERT: A Comprehensive Benchmark for Assessing Large Language Models' Safety through Red Teaming

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-12T14:13:30.244731Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:13:30.244731Z digest=sha256:c575060666b7e63ca91f1ff22a18fe08123562604ce4a8c7576a90276467b915

Observation de03f665-7dd2-4afe-8765-f0a6f20b74a9 · outbound

This paper cites Llama: Open and efficient foundation language models, 2023.

ChemSafetyBench: Benchmarking LLM Safety on Chemistry Domain Llama: Open and efficient foundation language models, 2023

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-12T14:13:30.273619Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:13:30.273619Z digest=sha256:97f3f383201fda197e085ac16d168657388909984added1ac7dc80312c4cce43

Observation fe1f287c-840b-4581-9fa9-a823de9a8de9 · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

ChemSafetyBench: Benchmarking LLM Safety on Chemistry Domain Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-12T14:13:30.280099Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:13:30.280099Z digest=sha256:9352b2fb2125ede9f0d5f774b6f9cc7b68c5ce82f55496e6f08950d4932c536d

Observation ae9b171d-8517-40e3-a7f8-8444b9f1f90f · outbound

This paper cites Decodingtrust: A comprehensive assessment of trustworthiness in gpt models.

ChemSafetyBench: Benchmarking LLM Safety on Chemistry Domain Decodingtrust: A comprehensive assessment of trustworthiness in gpt models

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:13:31.339126Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T14:13:30.286312Z digest=sha256:ada275ab85b19a9b534ea16beb6f8dceb2810889ff726e3b4ca3f8e895a9b376

Observation 6b0c52eb-18a4-4a0f-9cac-851a944fa2cb · outbound

This paper cites SciBench: Evaluating College-Level Scientific Problem-Solving Abilities of Large Language Models.

ChemSafetyBench: Benchmarking LLM Safety on Chemistry Domain SciBench: Evaluating College-Level Scientific Problem-Solving Abilities of Large Language Models

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-12T14:13:30.291357Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:13:30.291357Z digest=sha256:a238b5a54cbfad6461c42ac284614c0f6547e70c024cfefbf085a3916aa907b4

Observation a022faa0-2d92-4765-bce3-645f518d8ca1 · outbound

This paper cites Jailbroken: How does llm safety training fail? Advances in Neural Information Processing Systems , 36, 2024.

ChemSafetyBench: Benchmarking LLM Safety on Chemistry Domain Jailbroken: How does llm safety training fail? Advances in Neural Information Processing Systems , 36, 2024

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-12T14:13:30.314088Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:13:30.314088Z digest=sha256:2855eb2211dd198b5d78c1e48e4163549f5f048058a4e7c0f4d27e6e5809a059

Observation 9697a903-1d92-4c1f-9c92-73f4fefc761b · outbound

This paper cites Chemical weapons convention, 2024.

ChemSafetyBench: Benchmarking LLM Safety on Chemistry Domain Chemical weapons convention, 2024

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:13:31.307569Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T14:13:30.384200Z digest=sha256:62773ac329de19b3facc3896ab17b3432aed9c66ae199c04e145e0741cd49490

Observation db6229e4-2efd-48b2-97b1-9460259d5682 · outbound

This paper cites Controlled substances act, 2024.

ChemSafetyBench: Benchmarking LLM Safety on Chemistry Domain Controlled substances act, 2024

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:13:31.246001Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T14:13:30.464135Z digest=sha256:410d8701315c4cb81ffce8e0bff60c79fd07bb8706832f9f1c80835adb0f2095

Observation f577934a-aef6-4aab-9eda-6492d96d717c · outbound

This paper cites ReAct: Synergizing Reasoning and Acting in Language Models.

ChemSafetyBench: Benchmarking LLM Safety on Chemistry Domain ReAct: Synergizing Reasoning and Acting in Language Models

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-12T14:13:30.562189Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:13:30.562189Z digest=sha256:ef922fbd7a2d52a0d6a9844fb75bb2bdf486838aa9e71ae2b4ee395084453a65

Observation 734ff493-79bc-433a-8a94-aea1e4f23fec · outbound

This paper cites MARIO Eval: Evaluate Your Math LLM with your Math LLM--A mathematical dataset evaluation toolkit.

ChemSafetyBench: Benchmarking LLM Safety on Chemistry Domain MARIO Eval: Evaluate Your Math LLM with your Math LLM--A mathematical dataset evaluation toolkit

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-12T14:13:30.592858Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:13:30.592858Z digest=sha256:d0c920eaa417c3046a33a5b100c3ee48a443656ccd859e8a82cf97028392e84e

Observation 9e1e103c-3c95-44d3-a57a-c00c72da9f64 · outbound

This paper cites ChemLLM: A Chemical Large Language Model.

ChemSafetyBench: Benchmarking LLM Safety on Chemistry Domain ChemLLM: A Chemical Large Language Model

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-12T14:13:30.600439Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:13:30.600439Z digest=sha256:6367a1c70107e624b6f066fb5b6325dd4e517bcfc455e684334c906c0169ec84

Observation e0f023d7-c425-4bcc-83d2-415ce08046d3 · outbound

This paper cites JADE: A Linguistics-based Safety Evaluation Platform for Large Language Models.

ChemSafetyBench: Benchmarking LLM Safety on Chemistry Domain JADE: A Linguistics-based Safety Evaluation Platform for Large Language Models

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-12T14:13:30.606833Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:13:30.606833Z digest=sha256:45295ea199a5c180f77e06ead987c1389fd95fd123a5313378194cc5c54e1aab

Observation 484257d8-a543-48f8-8f52-60cdaa5fde61 · outbound

This paper cites SafetyBench: Evaluating the Safety of Large Language Models.

ChemSafetyBench: Benchmarking LLM Safety on Chemistry Domain SafetyBench: Evaluating the Safety of Large Language Models

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-12T14:13:30.612898Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:13:30.612898Z digest=sha256:fd7c61ee32892cea3a90e0b24b0e3edbb9c2080a81f349fc6c6b4c55958de473

Observation a0a33b7e-4953-4859-955f-a468d2ab28bb · outbound

This paper cites [[rating]].

ChemSafetyBench: Benchmarking LLM Safety on Chemistry Domain [[rating]]

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:13:31.129191Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T14:13:30.618840Z digest=sha256:72a4311ac22f7f5d5a0707b2521e8400ec288f02d3961342f1a5d1c6f204dd32

Pith citing papers

Observation 4185627b-7924-4961-90fa-c33d25ee590b · inbound

The Dual-use Dilemma in LLMs: Do Empowering Ethical Capacities Make a Degraded Utility? cites this paper.

The Dual-use Dilemma in LLMs: Do Empowering Ethical Capacities Make a Degraded Utility? ChemSafetyBench: Benchmarking LLM Safety on Chemistry Domain

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-10T18:31:03.512629Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T18:31:03.512629Z digest=sha256:94b0bd02c0a17bbe3e3eecd6a86ea9e76be9c27e7a6ad102cbe59ffc8952da05

Observation 4248f7c3-83cb-4f9f-b245-0cdec1a89edc · inbound

Bridging AI and Carbon Capture: A Dataset for LLMs in Ionic Liquids and CBE Research cites this paper.

Bridging AI and Carbon Capture: A Dataset for LLMs in Ionic Liquids and CBE Research ChemSafetyBench: Benchmarking LLM Safety on Chemistry Domain

Reference 108

Resolution
unresolved
no resolver link, observed 2026-08-15T22:31:28.922399Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T22:31:28.922399Z digest=sha256:d1ff5f4ab7026df2563e1befd47758e570e207134ee9c645700efae86caf7998

Observation cf9ff821-f42f-482a-8cc0-7a22da0f7e34 · inbound

Alignment and Safety in Large Language Models: Safety Mechanisms, Training Paradigms, and Emerging Challenges cites this paper.

Alignment and Safety in Large Language Models: Safety Mechanisms, Training Paradigms, and Emerging Challenges ChemSafetyBench: Benchmarking LLM Safety on Chemistry Domain

Reference 109

Resolution
unresolved
no resolver link, observed 2026-08-06T14:13:06.334084Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:13:06.334084Z digest=sha256:7e7ba44863389d04a4b151818efdaeccc7aca831d31b482e517afd3883c2076b

Observation 63b37aa9-321f-4d52-99f7-069b63dfb873 · inbound

Large Language Models Transform Organic Synthesis From Reaction Prediction to Automation cites this paper.

Large Language Models Transform Organic Synthesis From Reaction Prediction to Automation ChemSafetyBench: Benchmarking LLM Safety on Chemistry Domain

Reference 123

Resolution
unresolved
no resolver link, observed 2026-08-05T23:24:23.955660Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:24:23.955660Z digest=sha256:41ef384b2d9780a15eb6a8f490c6560726340686f7d430c6d1e103bc475e377b

Observation e93d3779-8b4b-46ec-a248-076e40f17437 · inbound

Boiling the Frog: A Multi-Turn Benchmark for Agentic Safety cites this paper.

Boiling the Frog: A Multi-Turn Benchmark for Agentic Safety ChemSafetyBench: Benchmarking LLM Safety on Chemistry Domain

Reference 96

Resolution
verified exact
arxiv_id, observed 2026-05-22T05:51:07.900654Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-22T05:50:28.114140Z digest=sha256:5ec98b20aff73c76fa005f53b03a7351c03d78a906528b8aa53fcf974e99fe52

Observation 495a9280-9279-4b0b-94be-9a4e07b58d48 · inbound

Boiling the Frog: A Multi-Turn Benchmark for Agentic Safety cites this paper.

Boiling the Frog: A Multi-Turn Benchmark for Agentic Safety ChemSafetyBench: Benchmarking LLM Safety on Chemistry Domain

Reference 96

Resolution
verified exact
arxiv_id, observed 2026-05-25T06:06:43.114555Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-25T06:05:27.736494Z digest=sha256:02881468a2897b1eff5c2ed6f64d57c40175602a1bbb413981b0b394c44e1a95

Observation aeafda4a-109f-4d1a-bd01-8ad5944fe80d · inbound

SciHazard: A Benchmark for Measuring Scientific Safety Risks with Decomposed Harm Scoring cites this paper.

SciHazard: A Benchmark for Measuring Scientific Safety Risks with Decomposed Harm Scoring ChemSafetyBench: Benchmarking LLM Safety on Chemistry Domain

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-01T14:45:43.513794Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T14:45:43.513794Z digest=sha256:511d158fee81f4c5b6b849ed2c55659dcba7c3a6b8821995ffe7d51d7b362ab7

Observation 8a6e282f-c5e9-493b-b647-a810e9a24c92 · inbound

onepot-Bench 0: towards lab-aware in silico chemistry benchmarks cites this paper.

onepot-Bench 0: towards lab-aware in silico chemistry benchmarks ChemSafetyBench: Benchmarking LLM Safety on Chemistry Domain

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-04T03:27:36.995699Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T03:27:36.995699Z digest=sha256:e9f3e8552be533d5a899cc626e2c14a89be8da570d7bd03f7a218031b3dbe9bd