Pith. sign in

Paper Citation Record · LEDGER

ChemSafetyBench: Benchmarking LLM Safety on Chemistry Domain

As of 18 August 2026, this Paper Citation Record lists 51 of 51 outbound references and 8 inbound Pith citation observations for arXiv:2411.16736.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2411.16736 v1

Coverage vector

measured 51 of 51 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T14:13:30.618840Z

measured 59 of 59 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 8 of 8 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T22:31:28.922399Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-25T06:06:43.111814Z

Reference resolution

51 of 51 outbound references displayed

  • verified exact0
  • verified fuzzy20
  • unresolved30
  • parse uncertain1
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 725413de-4f3b-47db-8a32-e6b56e87598d · outbound

This paper cites GPT-4 Technical Report.

ChemSafetyBench: Benchmarking LLM Safety on Chemistry Domain GPT-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-12T14:13:29.573144Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:13:29.573144Z digest=sha256:fecfe33a16046390b109defaf1c4625745a1b0b983226f92441ff9a49f78ed0a

Observation c73b7420-3bbe-481d-b3b7-334844a7032e · outbound

This paper cites an unresolved cited work.

ChemSafetyBench: Benchmarking LLM Safety on Chemistry Domain Unresolved cited work

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-12T14:13:29.597406Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:13:29.597406Z digest=sha256:d58217fa576908efae60009931d61c0f6db5c3ffec92f1d13fb578f5b373dc57

Observation e5cc7824-c022-4afc-9028-1391563cd134 · outbound

This paper cites Llama 3 model card.

ChemSafetyBench: Benchmarking LLM Safety on Chemistry Domain Llama 3 model card

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-12T14:13:29.604398Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:13:29.604398Z digest=sha256:62c083d7737481b41def64d2db3b6a93d7640819cd38fbc9e85c2c3199ea79f2

Observation 56c0175f-2563-4227-82d4-cf62b32ef8c8 · outbound

This paper cites The claude 3 model family: Opus, sonnet, haiku.

ChemSafetyBench: Benchmarking LLM Safety on Chemistry Domain The claude 3 model family: Opus, sonnet, haiku

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-12T14:13:29.610589Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:13:29.610589Z digest=sha256:f5f1f64f1930803657fbd005fc2f6d66a976376c3fba9e08d8353bacf04a7083

Observation 286ea39e-a229-4258-861a-65161cdf1422 · outbound

This paper cites Qwen Technical Report.

ChemSafetyBench: Benchmarking LLM Safety on Chemistry Domain Qwen Technical Report

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-12T14:13:29.633261Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:13:29.633261Z digest=sha256:677101a756ae82527f986feb827e69376624e959e00b5e1e8d82147a07211f43

Observation 6799db45-a564-423c-a4a5-04d43bcf6744 · outbound

This paper cites Autonomous chemical research with large language models.

ChemSafetyBench: Benchmarking LLM Safety on Chemistry Domain Autonomous chemical research with large language models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-12T14:13:29.674124Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:13:29.674124Z digest=sha256:a811d2f81f9811415f9957acc647d8d3cb153d7088015c68b276838acebd2f07

Observation 4b2ec2bb-5528-4b0f-8545-86056f5c2f64 · outbound

This paper cites Manning, and Roxana Daneshjou.

ChemSafetyBench: Benchmarking LLM Safety on Chemistry Domain Manning, and Roxana Daneshjou

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:13:32.656089Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T14:13:29.691608Z digest=sha256:f9514f5984a65f69b73ff76280d9487c4adc58938677a2443498f7b6a90223c2

Observation f2321ce5-37fc-4d13-8af0-31e70ba03428 · outbound

This paper cites Bioinfo-bench: A simple benchmark framework for llm bioinformatics skills evaluation.

ChemSafetyBench: Benchmarking LLM Safety on Chemistry Domain Bioinfo-bench: A simple benchmark framework for llm bioinformatics skills evaluation

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:13:32.436165Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T14:13:29.697690Z digest=sha256:eff6de021a35a4610fcfbb286cef8dbf706c2c0672c21486acaefbe9fcf2b0a9

Observation bc27f439-5fe0-421e-977d-03269b8582a1 · outbound

This paper cites 7 revealing ways ais fail: Neural networks can be disastrously brittle, forgetful, and surprisingly bad at math.

ChemSafetyBench: Benchmarking LLM Safety on Chemistry Domain 7 revealing ways ais fail: Neural networks can be disastrously brittle, forgetful, and surprisingly bad at math

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:13:32.407480Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T14:13:29.703002Z digest=sha256:8f9ef80242bcf9f286cd1d27ca9ce46f51d6badb12ba9550c4674854baec5318

Observation 92ce51a7-7600-4755-add1-36fd4f4e2828 · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

ChemSafetyBench: Benchmarking LLM Safety on Chemistry Domain Training Verifiers to Solve Math Word Problems

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-12T14:13:29.709737Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:13:29.709737Z digest=sha256:b13fa4a951f5706dd16cb776f082270924fbdf263e0047319d96ab0b1dd70194

Observation dcc1e499-e7a1-4f17-9e4a-aac80db084e5 · outbound

This paper cites Reaxys, 2024.

ChemSafetyBench: Benchmarking LLM Safety on Chemistry Domain Reaxys, 2024

Reference 11

Resolution
parse uncertain
raw_fallback, observed 2026-08-12T14:13:32.385695Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T14:13:29.716484Z digest=sha256:f0c64cd1c2ce4309cb43f088651e22a5d516e699f6643faa85b937e973da533e

Observation 603757da-0c17-470d-9cf5-5c377758aa20 · outbound

This paper cites Pubchem, 2024.

ChemSafetyBench: Benchmarking LLM Safety on Chemistry Domain Pubchem, 2024

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:13:32.361507Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T14:13:29.722191Z digest=sha256:0e4f399f0b963ddab94fd53896f52d2b79a3759c0541cff1bfa229d2d32974c6

Observation 862b633f-002e-46d1-a712-e0574b4be724 · outbound

This paper cites What can large language models do in chemistry? a comprehensive benchmark on eight tasks.

ChemSafetyBench: Benchmarking LLM Safety on Chemistry Domain What can large language models do in chemistry? a comprehensive benchmark on eight tasks

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:13:32.337335Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T14:13:29.753956Z digest=sha256:cd8b177eb44f4bb1a0c24bc9448516708f56c305bc7b44b620fe1435bf3c7c96

Observation 96573458-b7e6-437b-bbcf-409c957b5d8e · outbound

This paper cites Testing llm performance on the physics gre: some observations, 2023.

ChemSafetyBench: Benchmarking LLM Safety on Chemistry Domain Testing llm performance on the physics gre: some observations, 2023

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:13:32.313049Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T14:13:29.801405Z digest=sha256:b65178524b08320a265b881958365828d08739bb26e13bcdd8e3c98dd74d0297

Observation 2a72e99d-6cef-482d-b746-0ca9d31d536a · outbound

This paper cites A survey on large language models: Applications, challenges, limitations, and practical usage.

ChemSafetyBench: Benchmarking LLM Safety on Chemistry Domain A survey on large language models: Applications, challenges, limitations, and practical usage

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-12T14:13:29.806063Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:13:29.806063Z digest=sha256:e6fcbf4e9f47712e03317c27eb45b9796bf648320a8a9fad0d78ac0bfa4a1fb9

Observation 9ec6be15-c050-440e-96a4-388f71cc675b · outbound

This paper cites Control risk for potential misuse of artificial intelligence in science, 2023.

ChemSafetyBench: Benchmarking LLM Safety on Chemistry Domain Control risk for potential misuse of artificial intelligence in science, 2023

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:13:32.152244Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T14:13:29.810841Z digest=sha256:a64817f14cb08fe0d062ba2be82c4a25e74f4305860ad1baa7b91129b3df0dee

Observation b0f3e916-d60f-4496-a253-6d652f7d43a6 · outbound

This paper cites Measuring Mathematical Problem Solving With the MATH Dataset.

ChemSafetyBench: Benchmarking LLM Safety on Chemistry Domain Measuring Mathematical Problem Solving With the MATH Dataset

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-12T14:13:29.816346Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:13:29.816346Z digest=sha256:319b1a7dead3b28c519291d1ac3e847a40eafd2aa94714f6e4fc0d9d41cbc055

Observation 78ea54ef-4993-4424-9b62-327f83ec94b7 · outbound

This paper cites Amortizing intractable inference in large language models.

ChemSafetyBench: Benchmarking LLM Safety on Chemistry Domain Amortizing intractable inference in large language models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-12T14:13:29.822476Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:13:29.822476Z digest=sha256:c1ca43907a1da6dfe25b0cf6ce52657210da6f1ed013afb2c3f5f8c061e2f840

Observation dbebfee4-bf42-40d0-b12a-e4f429c54186 · outbound

This paper cites Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations.

ChemSafetyBench: Benchmarking LLM Safety on Chemistry Domain Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-12T14:13:29.831623Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:13:29.831623Z digest=sha256:16ba41fe9e38c20a17920c8a617c52edc7962fc60fede2e5a24d68476fddcea2

Observation 22ea5419-1a3e-4f19-a983-86722423f3c1 · outbound

This paper cites Leverag- ing large language models for predictive chemistry.

ChemSafetyBench: Benchmarking LLM Safety on Chemistry Domain Leverag- ing large language models for predictive chemistry

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:13:31.988696Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T14:13:29.837221Z digest=sha256:acffaa7c743f71a1c159281a4ebb414378db802091559376d2c41993b6dd17e0

Observation 372d6eb8-33bc-4048-931c-e419cb0d4501 · outbound

This paper cites A compre- hensive evaluation of large language models on benchmark biomedical text processing tasks.

ChemSafetyBench: Benchmarking LLM Safety on Chemistry Domain A compre- hensive evaluation of large language models on benchmark biomedical text processing tasks

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:13:31.950399Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T14:13:29.878291Z digest=sha256:c086063cd6e027343d801016a3af694a9a7cacf6235059956d384a58872d95b1

Observation c3a09434-197a-43db-8c09-1321c0593d69 · outbound

This paper cites an unresolved cited work.

ChemSafetyBench: Benchmarking LLM Safety on Chemistry Domain Unresolved cited work

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-12T14:13:29.912941Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:13:29.912941Z digest=sha256:c27af0c66f145c9ad7f7ea3082b5a6cbd3b5ea7be951bc0e79809462a0ca91cb

Observation 68a7c90c-a619-49eb-900f-684a2da1fcf4 · outbound

This paper cites Llm based biological named entity recognition from scientific literature.

ChemSafetyBench: Benchmarking LLM Safety on Chemistry Domain Llm based biological named entity recognition from scientific literature

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-12T14:13:29.919876Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:13:29.919876Z digest=sha256:c74c0065de263ad6737dec84ae1d226b9987911265a5bda9c3711e6262d5a270

Observation d0377570-475b-43ca-b87d-cb35a6b8dd0c · outbound

This paper cites Large Language Models in Plant Biology.

ChemSafetyBench: Benchmarking LLM Safety on Chemistry Domain Large Language Models in Plant Biology

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-12T14:13:29.928166Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:13:29.928166Z digest=sha256:721bc130ebb619371d182784dbd82d2931b1e310436c4b71d7bb4ed5e99d5e95

Observation 90d8e3c8-4f2f-4112-b717-4f6ae335f782 · outbound

This paper cites SALAD-Bench: A Hierarchical and Comprehensive Safety Benchmark for Large Language Models.

ChemSafetyBench: Benchmarking LLM Safety on Chemistry Domain SALAD-Bench: A Hierarchical and Comprehensive Safety Benchmark for Large Language Models

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-12T14:13:29.934618Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:13:29.934618Z digest=sha256:043ba790bd622e76b0d4b237e50ae20791a34f36e0adb18b4554fd4dc4800202

Observation 363eced4-2ff4-43d5-8778-2595877ebb02 · outbound

This paper cites Learn to explain: Multimodal reasoning via thought chains for science question answering.

ChemSafetyBench: Benchmarking LLM Safety on Chemistry Domain Learn to explain: Multimodal reasoning via thought chains for science question answering

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-12T14:13:29.941544Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:13:29.941544Z digest=sha256:4445002990a940671843acdb0908fdd9091c3aaceebf71b7a9339c93f8ed7364

Observation 283a0f19-1621-421c-98c2-33ccecf9eeef · outbound

This paper cites Bran, Sam Cox, Oliver Schilter, Carlo Baldassari, Andrew D White, and Philippe Schwaller.

ChemSafetyBench: Benchmarking LLM Safety on Chemistry Domain Bran, Sam Cox, Oliver Schilter, Carlo Baldassari, Andrew D White, and Philippe Schwaller

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:13:31.876810Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T14:13:29.950036Z digest=sha256:2a4f6f714290e8bdc6e2d328b76813b062710874aad6890ae201c3ddb42e703e

Observation 1286dc47-0650-45b9-b833-9a5a667c6282 · outbound

This paper cites Feasibility study on parameter adjustment for a humanoid using llm tailoring physical care.

ChemSafetyBench: Benchmarking LLM Safety on Chemistry Domain Feasibility study on parameter adjustment for a humanoid using llm tailoring physical care

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:13:31.815541Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T14:13:29.955973Z digest=sha256:d564160a2539d76d31f082246ef49bb1ff115fae85c64dea6c56b0054b8ba190

Observation aa4f2c8a-0342-449e-9f8c-483e2a8c807b · outbound

This paper cites Ghs classification results, 2024.

ChemSafetyBench: Benchmarking LLM Safety on Chemistry Domain Ghs classification results, 2024

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:13:31.730801Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T14:13:29.985121Z digest=sha256:cbe69563bb9a57a4e95de86f3e7d7c10dcf5c09da13858398b922e85bf1f7774

Observation 3b55005b-1c68-48cd-b89f-0506e4a0f26e · outbound

This paper cites Forbidden materials, 2024.

ChemSafetyBench: Benchmarking LLM Safety on Chemistry Domain Forbidden materials, 2024

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:13:31.655625Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T14:13:30.035469Z digest=sha256:23c349bafd0e632d09a1e40434c793f89cbf2454f4b3880e1aa5a64f5cb26fd2

Observation 954e25ab-33c8-4929-b8f1-380b590e7ecf · outbound

This paper cites SafetyPrompts: a Systematic Review of Open Datasets for Evaluating and Improving Large Language Model Safety.

ChemSafetyBench: Benchmarking LLM Safety on Chemistry Domain SafetyPrompts: a Systematic Review of Open Datasets for Evaluating and Improving Large Language Model Safety

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-12T14:13:30.092281Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:13:30.092281Z digest=sha256:8ea6169f96a9f2e7d61b27f7d637ecf6d88a574fbf766bf01de5d00773821384

Observation dafe91d1-599d-4f26-a3f6-5c41ef31160c · outbound

This paper cites LLM-Prop: Predicting Physical And Electronic Properties Of Crystalline Solids From Their Text Descriptions.

ChemSafetyBench: Benchmarking LLM Safety on Chemistry Domain LLM-Prop: Predicting Physical And Electronic Properties Of Crystalline Solids From Their Text Descriptions

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-12T14:13:30.100549Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:13:30.100549Z digest=sha256:7224068c3099d6ad04e773971e548d88a28be88b9f60096b4cb6ad881a0987d9

Observation 2b384e7b-4249-4d34-917a-c43584a04316 · outbound

This paper cites Sci- enceqa: A novel resource for question answering on scholarly articles.

ChemSafetyBench: Benchmarking LLM Safety on Chemistry Domain Sci- enceqa: A novel resource for question answering on scholarly articles

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-12T14:13:30.105680Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:13:30.105680Z digest=sha256:011b5199f10df45d17b2b8d319278379ca838f0832321f16e2f38a478237d071

Observation e8cebf45-4ba9-44e4-9a95-bca374ce03c5 · outbound

This paper cites Scifinder, 2024.

ChemSafetyBench: Benchmarking LLM Safety on Chemistry Domain Scifinder, 2024

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:13:31.476645Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T14:13:30.111553Z digest=sha256:db89c1a508da35452acd917f655e34e944d6c8afd608ddcff29031223ee8465e

Observation c4d34305-8f25-421a-a737-47ba038d9d36 · outbound

This paper cites Large language models encode clinical knowledge.

ChemSafetyBench: Benchmarking LLM Safety on Chemistry Domain Large language models encode clinical knowledge

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-12T14:13:30.116533Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:13:30.116533Z digest=sha256:e74ef7f713ca4426b4cda1cceb6fb2326f9e604e5e292fa43849f232c9f1e65a

Observation 3256f40b-4469-4936-8317-a2ed30ce3d90 · outbound

This paper cites Scieval: A multi-level large language model evaluation benchmark for scientific research, 2023.

ChemSafetyBench: Benchmarking LLM Safety on Chemistry Domain Scieval: A multi-level large language model evaluation benchmark for scientific research, 2023

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:13:31.390532Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T14:13:30.122858Z digest=sha256:95743088b48cf66fdfdda6f951f24905102e2ca2506919102662974432249383

Observation 29247775-4c1e-4820-bf7d-f694f3e14dfe · outbound

This paper cites Prioritizing safeguarding over autonomy: Risks of llm agents for science, 2024.

ChemSafetyBench: Benchmarking LLM Safety on Chemistry Domain Prioritizing safeguarding over autonomy: Risks of llm agents for science, 2024

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:13:31.371688Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T14:13:30.225702Z digest=sha256:ce6e2fb254e017dca636ee3e8d46a3aea83dacc9ed7c8fab023b1c91f97ac6d9

Observation 0acf68a5-75b1-4a76-8a61-9211234ece49 · outbound

This paper cites ALERT: A Comprehensive Benchmark for Assessing Large Language Models' Safety through Red Teaming.

ChemSafetyBench: Benchmarking LLM Safety on Chemistry Domain ALERT: A Comprehensive Benchmark for Assessing Large Language Models' Safety through Red Teaming

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-12T14:13:30.244731Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:13:30.244731Z digest=sha256:aeab1a388dac841bbb28da40be18c1bdc4cb4da272715a6e8374fb8c953f62a8

Observation de03f665-7dd2-4afe-8765-f0a6f20b74a9 · outbound

This paper cites Llama: Open and efficient foundation language models, 2023.

ChemSafetyBench: Benchmarking LLM Safety on Chemistry Domain Llama: Open and efficient foundation language models, 2023

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-12T14:13:30.273619Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:13:30.273619Z digest=sha256:356a06cf1c371c82e8d878a4f43833ae8dc424c74349bf9b9c6ab9dcf4acef27

Observation fe1f287c-840b-4581-9fa9-a823de9a8de9 · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

ChemSafetyBench: Benchmarking LLM Safety on Chemistry Domain Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-12T14:13:30.280099Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:13:30.280099Z digest=sha256:8d99b74bb34865ad261563c2c10d6b88b07429e178d87d480b5ae2f05f025830

Observation ae9b171d-8517-40e3-a7f8-8444b9f1f90f · outbound

This paper cites Decodingtrust: A comprehensive assessment of trustworthiness in gpt models.

ChemSafetyBench: Benchmarking LLM Safety on Chemistry Domain Decodingtrust: A comprehensive assessment of trustworthiness in gpt models

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:13:31.339126Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T14:13:30.286312Z digest=sha256:35561bb5dcc9f6d4a8bbce5bfe5438e639ce4d5619500ecdf152b909867a955a

Observation 6b0c52eb-18a4-4a0f-9cac-851a944fa2cb · outbound

This paper cites SciBench: Evaluating College-Level Scientific Problem-Solving Abilities of Large Language Models.

ChemSafetyBench: Benchmarking LLM Safety on Chemistry Domain SciBench: Evaluating College-Level Scientific Problem-Solving Abilities of Large Language Models

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-12T14:13:30.291357Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:13:30.291357Z digest=sha256:0eb3e23755447630ff9ba2b16fe7812edf2efe059acfe1649055131163a85131

Observation a022faa0-2d92-4765-bce3-645f518d8ca1 · outbound

This paper cites Jailbroken: How does llm safety training fail? Advances in Neural Information Processing Systems , 36, 2024.

ChemSafetyBench: Benchmarking LLM Safety on Chemistry Domain Jailbroken: How does llm safety training fail? Advances in Neural Information Processing Systems , 36, 2024

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-12T14:13:30.314088Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:13:30.314088Z digest=sha256:0cc536239e45e0f1dfce575a4f9dd01ef706233bdc12c020fa0a67a43ac6caf3

Observation 9697a903-1d92-4c1f-9c92-73f4fefc761b · outbound

This paper cites Chemical weapons convention, 2024.

ChemSafetyBench: Benchmarking LLM Safety on Chemistry Domain Chemical weapons convention, 2024

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:13:31.307569Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T14:13:30.384200Z digest=sha256:5454f242cd7231100e57c21dd18885e55a3580b3835f14218502207fb1334aeb

Observation db6229e4-2efd-48b2-97b1-9460259d5682 · outbound

This paper cites Controlled substances act, 2024.

ChemSafetyBench: Benchmarking LLM Safety on Chemistry Domain Controlled substances act, 2024

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:13:31.246001Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T14:13:30.464135Z digest=sha256:aae9d63fbbc248cc33c2ab31337714a093b024acd0cbcefd1a5f774a6bbcc104

Observation f577934a-aef6-4aab-9eda-6492d96d717c · outbound

This paper cites ReAct: Synergizing Reasoning and Acting in Language Models.

ChemSafetyBench: Benchmarking LLM Safety on Chemistry Domain ReAct: Synergizing Reasoning and Acting in Language Models

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-12T14:13:30.562189Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:13:30.562189Z digest=sha256:022f1f3fc94078fcd01a76c19373570160b621bcce1800291db3c90cb62fb5df

Observation 734ff493-79bc-433a-8a94-aea1e4f23fec · outbound

This paper cites MARIO Eval: Evaluate Your Math LLM with your Math LLM--A mathematical dataset evaluation toolkit.

ChemSafetyBench: Benchmarking LLM Safety on Chemistry Domain MARIO Eval: Evaluate Your Math LLM with your Math LLM--A mathematical dataset evaluation toolkit

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-12T14:13:30.592858Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:13:30.592858Z digest=sha256:e3450a60657c77a157c4e2047f9733b17eaee43f7610f7fa7ee5af3781fb03b7

Observation 9e1e103c-3c95-44d3-a57a-c00c72da9f64 · outbound

This paper cites ChemLLM: A Chemical Large Language Model.

ChemSafetyBench: Benchmarking LLM Safety on Chemistry Domain ChemLLM: A Chemical Large Language Model

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-12T14:13:30.600439Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:13:30.600439Z digest=sha256:edeef15ed8764be429725c18c0e10778c23b270258221b95ba5121f1e8a36293

Observation e0f023d7-c425-4bcc-83d2-415ce08046d3 · outbound

This paper cites JADE: A Linguistics-based Safety Evaluation Platform for Large Language Models.

ChemSafetyBench: Benchmarking LLM Safety on Chemistry Domain JADE: A Linguistics-based Safety Evaluation Platform for Large Language Models

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-12T14:13:30.606833Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:13:30.606833Z digest=sha256:1ad052132726146f2729acab3a04f45aed056f83012d5ff7573537e831e51614

Observation 484257d8-a543-48f8-8f52-60cdaa5fde61 · outbound

This paper cites SafetyBench: Evaluating the Safety of Large Language Models.

ChemSafetyBench: Benchmarking LLM Safety on Chemistry Domain SafetyBench: Evaluating the Safety of Large Language Models

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-12T14:13:30.612898Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:13:30.612898Z digest=sha256:ed6e55c14ba49c59f1c10ccf7a4d542344894612547530416dc162c5083c042e

Observation a0a33b7e-4953-4859-955f-a468d2ab28bb · outbound

This paper cites [[rating]].

ChemSafetyBench: Benchmarking LLM Safety on Chemistry Domain [[rating]]

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:13:31.129191Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T14:13:30.618840Z digest=sha256:6117edda7e64a837c30193ac3c1be00893f36ff2dd105b0864d2f88cb0d305b5

Pith citing papers

Observation 4185627b-7924-4961-90fa-c33d25ee590b · inbound

The Dual-use Dilemma in LLMs: Do Empowering Ethical Capacities Make a Degraded Utility? cites this paper.

The Dual-use Dilemma in LLMs: Do Empowering Ethical Capacities Make a Degraded Utility? ChemSafetyBench: Benchmarking LLM Safety on Chemistry Domain

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-10T18:31:03.512629Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T18:31:03.512629Z digest=sha256:d536ccf558f2b71202eda4f1fcaa96c7a6e0e26583094801afbdcf9db278d956

Observation 4248f7c3-83cb-4f9f-b245-0cdec1a89edc · inbound

Bridging AI and Carbon Capture: A Dataset for LLMs in Ionic Liquids and CBE Research cites this paper.

Bridging AI and Carbon Capture: A Dataset for LLMs in Ionic Liquids and CBE Research ChemSafetyBench: Benchmarking LLM Safety on Chemistry Domain

Reference 108

Resolution
unresolved
no resolver link, observed 2026-08-15T22:31:28.922399Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T22:31:28.922399Z digest=sha256:f0120a94d37a8a93b2804f7630bbd4ad0902d7b9784e76de29f79a8b3d14e655

Observation cf9ff821-f42f-482a-8cc0-7a22da0f7e34 · inbound

Alignment and Safety in Large Language Models: Safety Mechanisms, Training Paradigms, and Emerging Challenges cites this paper.

Alignment and Safety in Large Language Models: Safety Mechanisms, Training Paradigms, and Emerging Challenges ChemSafetyBench: Benchmarking LLM Safety on Chemistry Domain

Reference 109

Resolution
unresolved
no resolver link, observed 2026-08-06T14:13:06.334084Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:13:06.334084Z digest=sha256:4569b4532596c86618129092ffe6ed4fa401808dc8ab106efffde150c3879d55

Observation 63b37aa9-321f-4d52-99f7-069b63dfb873 · inbound

Large Language Models Transform Organic Synthesis From Reaction Prediction to Automation cites this paper.

Large Language Models Transform Organic Synthesis From Reaction Prediction to Automation ChemSafetyBench: Benchmarking LLM Safety on Chemistry Domain

Reference 123

Resolution
unresolved
no resolver link, observed 2026-08-05T23:24:23.955660Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:24:23.955660Z digest=sha256:f488cec47a7b96b3e53130adafcc86556e0d00c7167317c74a2b395040b8487d

Observation e93d3779-8b4b-46ec-a248-076e40f17437 · inbound

Boiling the Frog: A Multi-Turn Benchmark for Agentic Safety cites this paper.

Boiling the Frog: A Multi-Turn Benchmark for Agentic Safety ChemSafetyBench: Benchmarking LLM Safety on Chemistry Domain

Reference 96

Resolution
verified exact
arxiv_id, observed 2026-05-22T05:51:07.900654Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-22T05:50:28.114140Z digest=sha256:328f7f831967fa7223b227bf6995bedff6cac220fff96b10ec2595c77c2db970

Observation 495a9280-9279-4b0b-94be-9a4e07b58d48 · inbound

Boiling the Frog: A Multi-Turn Benchmark for Agentic Safety cites this paper.

Boiling the Frog: A Multi-Turn Benchmark for Agentic Safety ChemSafetyBench: Benchmarking LLM Safety on Chemistry Domain

Reference 96

Resolution
verified exact
arxiv_id, observed 2026-05-25T06:06:43.114555Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-25T06:05:27.736494Z digest=sha256:a2921a99b9b2010b0540c79d097b88f23d08ae2b814c4cffd8ee6d1f2cfe4f96

Observation aeafda4a-109f-4d1a-bd01-8ad5944fe80d · inbound

SciHazard: A Benchmark for Measuring Scientific Safety Risks with Decomposed Harm Scoring cites this paper.

SciHazard: A Benchmark for Measuring Scientific Safety Risks with Decomposed Harm Scoring ChemSafetyBench: Benchmarking LLM Safety on Chemistry Domain

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-01T14:45:43.513794Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T14:45:43.513794Z digest=sha256:5ea04d5744b3a429d829d34c97022bc2708c1d2601749e2cb0c045a7e13eb0ac

Observation 8a6e282f-c5e9-493b-b647-a810e9a24c92 · inbound

onepot-Bench 0: towards lab-aware in silico chemistry benchmarks cites this paper.

onepot-Bench 0: towards lab-aware in silico chemistry benchmarks ChemSafetyBench: Benchmarking LLM Safety on Chemistry Domain

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-04T03:27:36.995699Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T03:27:36.995699Z digest=sha256:5ffd77ad9cdfb1cc92e56552c5205d90d16a0602b99042055b0d6013b4758612