Pith. sign in

Paper Citation Record · LEDGER

RefusalBench: Why Refusal Rate Misranks Frontier LLMs on Biological Research Prompts

As of 18 August 2026, this Paper Citation Record lists 41 of 41 outbound references and 2 inbound Pith citation observations for arXiv:2605.21545.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2605.21545 v1

Coverage vector

measured 41 of 41 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-22T01:19:00.268857Z

measured 43 of 43 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-02T08:40:01.438522Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

41 of 41 outbound references displayed

  • verified exact17
  • verified fuzzy19
  • unresolved2
  • parse uncertain0
  • malformed identifier3
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation ca587e6c-8cdc-440a-8aeb-92d5ed1b749a · outbound

This paper cites One-shot design of functional protein binders with BindCraft.

RefusalBench: Why Refusal Rate Misranks Frontier LLMs on Biological Research Prompts One-shot design of functional protein binders with BindCraft

Reference 1

Resolution
verified exact
doi, observed 2026-05-22T01:20:51.616120Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-22T01:19:00.268857Z digest=sha256:4d6b82a4f27207690994e24fe1c312c2e624beca946fc7752fc2291c74fe28f1

Observation 87ea9d68-9a98-481e-931a-9a192853a492 · outbound

This paper cites ProteinCrow: A language model agent that can design proteins.

RefusalBench: Why Refusal Rate Misranks Frontier LLMs on Biological Research Prompts ProteinCrow: A language model agent that can design proteins

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T01:20:52.359382Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-22T01:19:00.268857Z digest=sha256:63f90b33c60c458a16d350720aa97230c92538b95303c35a3d32ed58b9fdb1d2

Observation 71322988-55b6-48c7-8f07-74e2ec7021db · outbound

This paper cites an unresolved cited work.

RefusalBench: Why Refusal Rate Misranks Frontier LLMs on Biological Research Prompts Unresolved cited work

Reference 3

Resolution
unresolved
raw_fallback, observed 2026-05-22T01:20:52.362821Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-22T01:19:00.268857Z digest=sha256:fe3bafe8420b369b4a1f36e63f0e33b015cf672940b9fa6b9b586e90387dd91b

Observation c5d25555-0c66-4abb-b731-6156427807f6 · outbound

This paper cites Beyond protein language models: An agentic LLM framework for mechanistic enzyme design.

RefusalBench: Why Refusal Rate Misranks Frontier LLMs on Biological Research Prompts Beyond protein language models: An agentic LLM framework for mechanistic enzyme design

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T01:20:52.366248Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-22T01:19:00.268857Z digest=sha256:1f03f7f83dcdafc7a87305ff55c4ace5910353bb13a4c793edbb1c22ddd5d248

Observation 8cdd4475-adb8-433d-bf8c-f5138fa06e56 · outbound

This paper cites arXiv:2511.19423.

RefusalBench: Why Refusal Rate Misranks Frontier LLMs on Biological Research Prompts arXiv:2511.19423

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-22T01:20:51.962759Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-22T01:19:00.268857Z digest=sha256:408d097e49248fb49308ca78c5235dd2b567011efe84dc8a389c27e6ef6e387e

Observation 84ecc782-35d6-44ae-803c-60d44eb0e743 · outbound

This paper cites ProtoCy- cle: Reflective tool-augmented planning for text-guided protein design.

RefusalBench: Why Refusal Rate Misranks Frontier LLMs on Biological Research Prompts ProtoCy- cle: Reflective tool-augmented planning for text-guided protein design

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T01:20:52.337718Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-22T01:19:00.268857Z digest=sha256:1a1f063170e33b2a09fb26d324a4741b44598c6fa72561c198f72ebef6395ba4

Observation 8c811739-6118-49d1-adad-19f4f0a926db · outbound

This paper cites ProtoCycle: Reflective Tool-Augmented Planning for Text-Guided Protein Design.

RefusalBench: Why Refusal Rate Misranks Frontier LLMs on Biological Research Prompts ProtoCycle: Reflective Tool-Augmented Planning for Text-Guided Protein Design

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-05-22T01:20:51.957519Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-22T01:19:00.268857Z digest=sha256:c3819fcc547107d84ac0f6bb3e8f4bead18a250810ee9d993918677b67dd3dfc

Observation 1ec75940-ec6f-40b1-84a5-add990c4bbc5 · outbound

This paper cites ProteinMCP: An agentic AI framework for autonomous protein engineer- ing.

RefusalBench: Why Refusal Rate Misranks Frontier LLMs on Biological Research Prompts ProteinMCP: An agentic AI framework for autonomous protein engineer- ing

Reference 8

Resolution
verified exact
doi, observed 2026-05-22T01:20:51.596971Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-22T01:19:00.268857Z digest=sha256:c811092a61c4dfe9efb6288206af60c95de008e994d69fdcdcbc18b21ada500d

Observation 1af6843a-8940-415d-b357-703c74c64aeb · outbound

This paper cites Using a GPT-5-driven autonomous lab to optimize the cost and titer of cell-free protein synthesis.

RefusalBench: Why Refusal Rate Misranks Frontier LLMs on Biological Research Prompts Using a GPT-5-driven autonomous lab to optimize the cost and titer of cell-free protein synthesis

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T01:20:52.333877Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-22T01:19:00.268857Z digest=sha256:191c9c04cbc2b4f99f62aaf597675100cd39fcdcd6156bdf6023d0c54a315b32

Observation 59b0a91d-7833-4bb1-b207-2e7a3cfbef34 · outbound

This paper cites an unresolved cited work.

RefusalBench: Why Refusal Rate Misranks Frontier LLMs on Biological Research Prompts Unresolved cited work

Reference 10

Resolution
verified exact
doi, observed 2026-05-22T01:20:51.583014Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-22T01:19:00.268857Z digest=sha256:6f92d1486b1110c19a6d951ca41398ae533f0f34427444532bd0387320e54284

Observation ea3fd858-7d1d-49a2-9ad6-0e99a0684465 · outbound

This paper cites Agen- tic BAIM–LLM evaluation (ABLE): Bench- marking LLM use of protein design tools.

RefusalBench: Why Refusal Rate Misranks Frontier LLMs on Biological Research Prompts Agen- tic BAIM–LLM evaluation (ABLE): Bench- marking LLM use of protein design tools

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T01:20:52.297123Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-22T01:19:00.268857Z digest=sha256:65f4920a08aaec510e6d3841847e9497092fc24796b6daa067c495a2cb3e669f

Observation 785c5b59-3e0a-4d58-af23-a26f307823ed · outbound

This paper cites XSTest: A Test Suite for Identifying Exaggerated Safety Behaviours in Large Language Models.

RefusalBench: Why Refusal Rate Misranks Frontier LLMs on Biological Research Prompts XSTest: A Test Suite for Identifying Exaggerated Safety Behaviours in Large Language Models

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-05-22T01:20:51.901268Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-22T01:19:00.268857Z digest=sha256:5cce5fab8728ddcb24a6dfcf7ac88f356045eca6cde6717d0d56d9c1fa0515e0

Observation 7c5153cb-46d4-4fde-a475-6752eb7f8bf2 · outbound

This paper cites OR-Bench: An over- refusal benchmark for large language mod- els.

RefusalBench: Why Refusal Rate Misranks Frontier LLMs on Biological Research Prompts OR-Bench: An over- refusal benchmark for large language mod- els

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T01:20:52.305301Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-22T01:19:00.268857Z digest=sha256:72da173bd00143ef5b4036f4796b14220880a56f7fb1944f9814bde5ca0f07b0

Observation 757b7704-4f27-4c1d-9dd3-cfca173d462b · outbound

This paper cites OR-Bench: An Over-Refusal Benchmark for Large Language Models.

RefusalBench: Why Refusal Rate Misranks Frontier LLMs on Biological Research Prompts OR-Bench: An Over-Refusal Benchmark for Large Language Models

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-22T01:20:51.889010Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-22T01:19:00.268857Z digest=sha256:87e9b9d35ff563ae6b223c8a827f4bd9682d68c615ea645a240072368a8c7987

Observation fded0e28-e891-4b7a-8aff-bed07bfe1418 · outbound

This paper cites Cannot or Should Not? Automatic Analysis of Refusal Composition in IFT/RLHF Datasets and Refusal Behavior of Black-Box LLMs.

RefusalBench: Why Refusal Rate Misranks Frontier LLMs on Biological Research Prompts Cannot or Should Not? Automatic Analysis of Refusal Composition in IFT/RLHF Datasets and Refusal Behavior of Black-Box LLMs

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-22T01:20:51.947425Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-22T01:19:00.268857Z digest=sha256:df946232ded11e08896d6d4efdb7004eed1d79e5228e96154826980625450669

Observation 33cd4966-0a11-4cb1-9c44-6384baee4b01 · outbound

This paper cites Forbid- den science: Dual-use AI challenge benchmark and scientific refusal tests.

RefusalBench: Why Refusal Rate Misranks Frontier LLMs on Biological Research Prompts Forbid- den science: Dual-use AI challenge benchmark and scientific refusal tests

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T01:20:52.330044Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-22T01:19:00.268857Z digest=sha256:ee4c1cd41b3d9043bf3e05f7a5386268af3937549d1c79b165ea390fde9f4b06

Observation 701f45cb-5f99-4851-b73d-42b3d9100d30 · outbound

This paper cites Forbidden Science: Dual-Use AI Challenge Benchmark and Scientific Refusal Tests.

RefusalBench: Why Refusal Rate Misranks Frontier LLMs on Biological Research Prompts Forbidden Science: Dual-Use AI Challenge Benchmark and Scientific Refusal Tests

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-22T01:20:51.936763Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-22T01:19:00.268857Z digest=sha256:2fc598185c77c2ed70a9e0f2ba76c2da59b23f1c65450fb1ccbb9dc82015b11e

Observation 7deeae4d-88e2-4b7c-ba1b-3b23228cd846 · outbound

This paper cites Political censorship in large language models originating from China.

RefusalBench: Why Refusal Rate Misranks Frontier LLMs on Biological Research Prompts Political censorship in large language models originating from China

Reference 18

Resolution
malformed identifier
raw_fallback, observed 2026-05-22T01:20:52.313076Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-22T01:19:00.268857Z digest=sha256:1e55efeaf74c4ecb73791594b6d5790ed2881797f8edc892b792425a091b8667

Observation 40a6b9f7-813d-4825-b7c4-f485731ec9a6 · outbound

This paper cites Virol- ogy capabilities test (VCT): A multimodal virology Q&A benchmark.

RefusalBench: Why Refusal Rate Misranks Frontier LLMs on Biological Research Prompts Virol- ogy capabilities test (VCT): A multimodal virology Q&A benchmark

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T01:20:52.326050Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-22T01:19:00.268857Z digest=sha256:e449163c92c1a77be217d1a9aa664a532aaa19e5cc4528ba5884b940f3c23732

Observation 9a754cc8-6fba-471a-8dfe-81bb214e9066 · outbound

This paper cites Virology Capabilities Test (VCT): A Multimodal Virology Q&A Benchmark.

RefusalBench: Why Refusal Rate Misranks Frontier LLMs on Biological Research Prompts Virology Capabilities Test (VCT): A Multimodal Virology Q&A Benchmark

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-22T01:20:51.895253Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-22T01:19:00.268857Z digest=sha256:e6af159d2cf8e7e5d9e6db4d7e8b64cd670605d01f1c08d078f8f7ad31339b34

Observation aeb64051-00e6-4474-9ae2-1a7f74900b65 · outbound

This paper cites Can large language models democratize access to dual-use biotechnology? arXiv preprint.

RefusalBench: Why Refusal Rate Misranks Frontier LLMs on Biological Research Prompts Can large language models democratize access to dual-use biotechnology? arXiv preprint

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T01:20:52.369823Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-22T01:19:00.268857Z digest=sha256:cb37936f83826b0113d63638780c9c0a6b8021f80c7a9cc2e73d55c3d3e20256

Observation 9c3e08c0-dcf5-4aa6-b0e2-719652a1dbf0 · outbound

This paper cites Can large language models democratize access to dual-use biotechnology?.

RefusalBench: Why Refusal Rate Misranks Frontier LLMs on Biological Research Prompts Can large language models democratize access to dual-use biotechnology?

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-22T01:20:51.931094Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-22T01:19:00.268857Z digest=sha256:8a86e227f8a3120e4ca1e7eef8ab09b3d80bba13ba4d78ec55da23713161e4a6

Observation 647aa87d-09ae-4956-a3ea-6f1ace9337b1 · outbound

This paper cites The next-generation Open Targets platform: reimag- ined, redesigned, rebuilt.

RefusalBench: Why Refusal Rate Misranks Frontier LLMs on Biological Research Prompts The next-generation Open Targets platform: reimag- ined, redesigned, rebuilt

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T01:20:52.348456Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-22T01:19:00.268857Z digest=sha256:95e3086f6730b1f7f0184de7d301ecf0fe97b26be5dd6433f4d0abf626e4da40

Observation 107763b1-d770-4f69-8057-3f1d10f4d527 · outbound

This paper cites UniProt: the uni- versal protein knowledgebase in 2025.

RefusalBench: Why Refusal Rate Misranks Frontier LLMs on Biological Research Prompts UniProt: the uni- versal protein knowledgebase in 2025

Reference 24

Resolution
verified exact
doi, observed 2026-05-22T01:20:51.610655Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-22T01:19:00.268857Z digest=sha256:f555a906396c04f6fd3415f1d114315cc097282140d5017badbdf16205028f72

Observation bc2c5a7a-4cc5-4e1a-aa5a-85986ca94a4e · outbound

This paper cites NVIDIA Nemotron 3 Su- per: A 120B hybrid Mamba-Transformer MoE model for agentic reasoning.

RefusalBench: Why Refusal Rate Misranks Frontier LLMs on Biological Research Prompts NVIDIA Nemotron 3 Su- per: A 120B hybrid Mamba-Transformer MoE model for agentic reasoning

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T01:20:52.320830Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-22T01:19:00.268857Z digest=sha256:6fd867aa047717efde6b9319e5536f1138aa281be1cd7e97d9016a57d2e3c9e6

Observation a75acdab-0de0-420b-8df7-c04394647de6 · outbound

This paper cites Open-weights release; 12B active / 120B total parameters; 1M-token context.

RefusalBench: Why Refusal Rate Misranks Frontier LLMs on Biological Research Prompts Open-weights release; 12B active / 120B total parameters; 1M-token context

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T01:20:52.301203Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-22T01:19:00.268857Z digest=sha256:2f907b86fc39ce505b41af6f9abf63ae2a34f55dde64028f45f4c1c319aa6682

Observation c6013dec-24c1-446d-9c90-83ac80a86860 · outbound

This paper cites SORRY- Bench: Systematically evaluating large lan- guage model safety refusal.

RefusalBench: Why Refusal Rate Misranks Frontier LLMs on Biological Research Prompts SORRY- Bench: Systematically evaluating large lan- guage model safety refusal

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T01:20:52.316971Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-22T01:19:00.268857Z digest=sha256:63838bedc1d34ef5f9c6d9de770143bc9eacc776f5e40013d1e3404f11c62324

Observation 6b539532-c25e-4445-ad80-74ef7a44f0c4 · outbound

This paper cites SORRY-Bench: Systematically Evaluating Large Language Model Safety Refusal.

RefusalBench: Why Refusal Rate Misranks Frontier LLMs on Biological Research Prompts SORRY-Bench: Systematically Evaluating Large Language Model Safety Refusal

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-05-22T01:20:51.907305Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-22T01:19:00.268857Z digest=sha256:1f6522fc5c2d3f8ec74cabe9a57eb0e1fa1660371394222608c9830555d12a60

Observation ffa28ec7-3751-42ba-a964-4e66f42c28d6 · outbound

This paper cites Content Analysis: An Introduction to Its Methodology.

RefusalBench: Why Refusal Rate Misranks Frontier LLMs on Biological Research Prompts Content Analysis: An Introduction to Its Methodology

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T01:20:52.352023Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-22T01:19:00.268857Z digest=sha256:9dc3efcf231d4760bb526ce3f93b4141c153b51c527a570fe90749e278aa8a16

Observation 00fff7df-d088-4920-ab28-8219d46a02cc · outbound

This paper cites The art of saying no: Contextual non- compliance in language models.

RefusalBench: Why Refusal Rate Misranks Frontier LLMs on Biological Research Prompts The art of saying no: Contextual non- compliance in language models

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T01:20:52.344699Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-22T01:19:00.268857Z digest=sha256:174f40e4719369c6503bb88150026a84070470009af45c6bd7bf0fe7db82b39b

Observation 1b504a5b-3776-4a91-9246-4843a4b2df69 · outbound

This paper cites The Art of Saying No: Contextual Noncompliance in Language Models.

RefusalBench: Why Refusal Rate Misranks Frontier LLMs on Biological Research Prompts The Art of Saying No: Contextual Noncompliance in Language Models

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-05-22T01:20:51.941888Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-22T01:19:00.268857Z digest=sha256:db6a82e54ee0653d56ccc66ef8882e6cd54f3f84d3bda44600a06e9fef40312e

Observation b5d7b3ba-568f-4d42-aa10-7fa20ca245b4 · outbound

This paper cites Constitutional AI: Harmlessness from AI Feedback.

RefusalBench: Why Refusal Rate Misranks Frontier LLMs on Biological Research Prompts Constitutional AI: Harmlessness from AI Feedback

Reference 32

Resolution
verified exact
local_arxiv, observed 2026-05-22T01:20:51.919483Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-22T01:19:00.268857Z digest=sha256:9e88726311a5a29fcd7cdee0585b805818304b24b9c4b0dcd0a15200ee72e712

Observation cd4dd58d-1202-454f-a890-7ae279186d4c · outbound

This paper cites Anthropic’s responsible scaling policy.

RefusalBench: Why Refusal Rate Misranks Frontier LLMs on Biological Research Prompts Anthropic’s responsible scaling policy

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T01:20:52.341138Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-22T01:19:00.268857Z digest=sha256:3049382d530819c25f42aa4378ecac93b1c231540dda6d5539f5333db67a7e8f

Observation 6d23fe70-0226-4931-91fe-1975d52042fc · outbound

This paper cites A., Mathur, S., Salabert, D., Ballot, J., R´egulo, C., Metcalfe, T.

RefusalBench: Why Refusal Rate Misranks Frontier LLMs on Biological Research Prompts A., Mathur, S., Salabert, D., Ballot, J., R´egulo, C., Metcalfe, T

Reference 34

Resolution
malformed identifier
doi_truncated, observed 2026-05-22T01:20:51.604548Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-22T01:19:00.268857Z digest=sha256:7b40e98fa4b4ffa622483628188d132936e2ad7d3fc18f24e118752601a110cc

Observation da7d68e0-3579-4211-b94c-d698de212477 · outbound

This paper cites Dual use of ar- tificial intelligence-powered drug discovery.

RefusalBench: Why Refusal Rate Misranks Frontier LLMs on Biological Research Prompts Dual use of ar- tificial intelligence-powered drug discovery

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T01:20:52.373431Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-22T01:19:00.268857Z digest=sha256:70420092a4f40b88d7bcf0334b7b44389e327817ccc2bd39f22fa467fb6c8139

Observation 4bce7a55-a6b8-43ed-814e-3a9912cdd7a3 · outbound

This paper cites an unresolved cited work.

RefusalBench: Why Refusal Rate Misranks Frontier LLMs on Biological Research Prompts Unresolved cited work

Reference 36

Resolution
unresolved
raw_fallback, observed 2026-05-22T01:20:52.355717Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-22T01:19:00.268857Z digest=sha256:41b1bb3990147e9b07ff7dccebc6a6d52bfb376c313aa6c808635de4a57f8611

Observation a5dcc4fd-ec7f-48e4-9b1a-0b0cac51579a · outbound

This paper cites Training language models to follow instructions with hu- man feedback.

RefusalBench: Why Refusal Rate Misranks Frontier LLMs on Biological Research Prompts Training language models to follow instructions with hu- man feedback

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T01:20:52.309141Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-22T01:19:00.268857Z digest=sha256:fffe8b299789d69d2be1f7bafc47c55a02d33af93a44fbe8ca780e08223eca72

Observation a9b5081e-678c-4a61-bda6-d4b89781c5c7 · outbound

This paper cites Training language models to follow instructions with human feedback.

RefusalBench: Why Refusal Rate Misranks Frontier LLMs on Biological Research Prompts Training language models to follow instructions with human feedback

Reference 38

Resolution
verified exact
local_arxiv, observed 2026-05-22T01:20:51.952881Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-22T01:19:00.268857Z digest=sha256:487e5d3140c4a6c428f7665c26adaaee50376d41badb790b983358bb86b26b27

Observation a9a33a66-c8c8-4fb4-be63-55d036d2914b · outbound

This paper cites Harm- Bench: A standardized evaluation framework for automated red teaming and robust re- fusal.

RefusalBench: Why Refusal Rate Misranks Frontier LLMs on Biological Research Prompts Harm- Bench: A standardized evaluation framework for automated red teaming and robust re- fusal

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T01:20:52.376614Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-22T01:19:00.268857Z digest=sha256:706995cd84eab72aad88b58b20685768f80bb1352a659293d6c12d861a9e5211

Observation 07c00f47-c693-4eea-96ce-da6e2eed708c · outbound

This paper cites HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust Refusal.

RefusalBench: Why Refusal Rate Misranks Frontier LLMs on Biological Research Prompts HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust Refusal

Reference 40

Resolution
verified exact
local_arxiv, observed 2026-05-22T01:20:51.925495Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-22T01:19:00.268857Z digest=sha256:90db648715b5f7d7a1fec5944e6324c7e32c404d1a3754faed704bf1a99b93c2

Observation a356293f-7cbf-43d2-9107-6d4e18597a37 · outbound

This paper cites The WMDP Benchmark: Measuring and Reducing Malicious Use With Unlearning.

RefusalBench: Why Refusal Rate Misranks Frontier LLMs on Biological Research Prompts The WMDP Benchmark: Measuring and Reducing Malicious Use With Unlearning

Reference 41

Resolution
malformed identifier
local_arxiv, observed 2026-05-22T01:20:51.913107Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-22T01:19:00.268857Z digest=sha256:9027315d94b33d18a1aa62283be723bfff0c1a491118274532c0b7b693980f11

Pith citing papers

Observation 910685a1-f3f7-4069-97bd-72855cabab83 · inbound

BioSecBench-Refusal: A paired metric for performance and alignment in agentic biosecurity risk assessment cites this paper.

BioSecBench-Refusal: A paired metric for performance and alignment in agentic biosecurity risk assessment RefusalBench: Why Refusal Rate Misranks Frontier LLMs on Biological Research Prompts

Reference 12

Resolution
unresolved
no resolver link, observed 2026-07-11T16:42:49.284801Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T16:42:49.284801Z digest=sha256:d8eb0a8c9be285ab4c619685f61a978012861d1e3f12926687f38de0fb8f4040

Observation 9e8cf392-3063-4472-94f6-bd8b27376ef6 · inbound

BioSecBench-Refusal: A paired metric for performance and alignment in agentic biosecurity risk assessment cites this paper.

BioSecBench-Refusal: A paired metric for performance and alignment in agentic biosecurity risk assessment RefusalBench: Why Refusal Rate Misranks Frontier LLMs on Biological Research Prompts

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-02T08:40:01.438522Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T08:40:01.438522Z digest=sha256:8bd94b9d51cbb030f1290095652746de57447f05a9b0c4948d785c2ed9ffd4ae