Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-22T01:19:00.268857Z
Paper Citation Record · LEDGER
As of 18 August 2026, this Paper Citation Record lists 41 of 41 outbound references and 2 inbound Pith citation observations for arXiv:2605.21545.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-22T01:19:00.268857Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-02T08:40:01.438522Z
A source-named dated measurement, never combined with another source.
Source: cited_works
41 of 41 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation ca587e6c-8cdc-440a-8aeb-92d5ed1b749a · outbound
RefusalBench: Why Refusal Rate Misranks Frontier LLMs on Biological Research Prompts One-shot design of functional protein binders with BindCraft
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 87ea9d68-9a98-481e-931a-9a192853a492 · outbound
RefusalBench: Why Refusal Rate Misranks Frontier LLMs on Biological Research Prompts ProteinCrow: A language model agent that can design proteins
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 71322988-55b6-48c7-8f07-74e2ec7021db · outbound
RefusalBench: Why Refusal Rate Misranks Frontier LLMs on Biological Research Prompts Unresolved cited work
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation c5d25555-0c66-4abb-b731-6156427807f6 · outbound
RefusalBench: Why Refusal Rate Misranks Frontier LLMs on Biological Research Prompts Beyond protein language models: An agentic LLM framework for mechanistic enzyme design
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 8cdd4475-adb8-433d-bf8c-f5138fa06e56 · outbound
RefusalBench: Why Refusal Rate Misranks Frontier LLMs on Biological Research Prompts arXiv:2511.19423
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 84ecc782-35d6-44ae-803c-60d44eb0e743 · outbound
RefusalBench: Why Refusal Rate Misranks Frontier LLMs on Biological Research Prompts ProtoCy- cle: Reflective tool-augmented planning for text-guided protein design
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 8c811739-6118-49d1-adad-19f4f0a926db · outbound
RefusalBench: Why Refusal Rate Misranks Frontier LLMs on Biological Research Prompts ProtoCycle: Reflective Tool-Augmented Planning for Text-Guided Protein Design
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 1ec75940-ec6f-40b1-84a5-add990c4bbc5 · outbound
RefusalBench: Why Refusal Rate Misranks Frontier LLMs on Biological Research Prompts ProteinMCP: An agentic AI framework for autonomous protein engineer- ing
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 1af6843a-8940-415d-b357-703c74c64aeb · outbound
RefusalBench: Why Refusal Rate Misranks Frontier LLMs on Biological Research Prompts Using a GPT-5-driven autonomous lab to optimize the cost and titer of cell-free protein synthesis
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 59b0a91d-7833-4bb1-b207-2e7a3cfbef34 · outbound
RefusalBench: Why Refusal Rate Misranks Frontier LLMs on Biological Research Prompts Unresolved cited work
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation ea3fd858-7d1d-49a2-9ad6-0e99a0684465 · outbound
RefusalBench: Why Refusal Rate Misranks Frontier LLMs on Biological Research Prompts Agen- tic BAIM–LLM evaluation (ABLE): Bench- marking LLM use of protein design tools
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 785c5b59-3e0a-4d58-af23-a26f307823ed · outbound
RefusalBench: Why Refusal Rate Misranks Frontier LLMs on Biological Research Prompts XSTest: A Test Suite for Identifying Exaggerated Safety Behaviours in Large Language Models
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 7c5153cb-46d4-4fde-a475-6752eb7f8bf2 · outbound
RefusalBench: Why Refusal Rate Misranks Frontier LLMs on Biological Research Prompts OR-Bench: An over- refusal benchmark for large language mod- els
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 757b7704-4f27-4c1d-9dd3-cfca173d462b · outbound
RefusalBench: Why Refusal Rate Misranks Frontier LLMs on Biological Research Prompts OR-Bench: An Over-Refusal Benchmark for Large Language Models
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation fded0e28-e891-4b7a-8aff-bed07bfe1418 · outbound
RefusalBench: Why Refusal Rate Misranks Frontier LLMs on Biological Research Prompts Cannot or Should Not? Automatic Analysis of Refusal Composition in IFT/RLHF Datasets and Refusal Behavior of Black-Box LLMs
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 33cd4966-0a11-4cb1-9c44-6384baee4b01 · outbound
RefusalBench: Why Refusal Rate Misranks Frontier LLMs on Biological Research Prompts Forbid- den science: Dual-use AI challenge benchmark and scientific refusal tests
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 701f45cb-5f99-4851-b73d-42b3d9100d30 · outbound
RefusalBench: Why Refusal Rate Misranks Frontier LLMs on Biological Research Prompts Forbidden Science: Dual-Use AI Challenge Benchmark and Scientific Refusal Tests
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 7deeae4d-88e2-4b7c-ba1b-3b23228cd846 · outbound
RefusalBench: Why Refusal Rate Misranks Frontier LLMs on Biological Research Prompts Political censorship in large language models originating from China
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 40a6b9f7-813d-4825-b7c4-f485731ec9a6 · outbound
RefusalBench: Why Refusal Rate Misranks Frontier LLMs on Biological Research Prompts Virol- ogy capabilities test (VCT): A multimodal virology Q&A benchmark
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 9a754cc8-6fba-471a-8dfe-81bb214e9066 · outbound
RefusalBench: Why Refusal Rate Misranks Frontier LLMs on Biological Research Prompts Virology Capabilities Test (VCT): A Multimodal Virology Q&A Benchmark
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation aeb64051-00e6-4474-9ae2-1a7f74900b65 · outbound
RefusalBench: Why Refusal Rate Misranks Frontier LLMs on Biological Research Prompts Can large language models democratize access to dual-use biotechnology? arXiv preprint
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 9c3e08c0-dcf5-4aa6-b0e2-719652a1dbf0 · outbound
RefusalBench: Why Refusal Rate Misranks Frontier LLMs on Biological Research Prompts Can large language models democratize access to dual-use biotechnology?
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 647aa87d-09ae-4956-a3ea-6f1ace9337b1 · outbound
RefusalBench: Why Refusal Rate Misranks Frontier LLMs on Biological Research Prompts The next-generation Open Targets platform: reimag- ined, redesigned, rebuilt
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 107763b1-d770-4f69-8057-3f1d10f4d527 · outbound
RefusalBench: Why Refusal Rate Misranks Frontier LLMs on Biological Research Prompts UniProt: the uni- versal protein knowledgebase in 2025
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation bc2c5a7a-4cc5-4e1a-aa5a-85986ca94a4e · outbound
RefusalBench: Why Refusal Rate Misranks Frontier LLMs on Biological Research Prompts NVIDIA Nemotron 3 Su- per: A 120B hybrid Mamba-Transformer MoE model for agentic reasoning
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation a75acdab-0de0-420b-8df7-c04394647de6 · outbound
RefusalBench: Why Refusal Rate Misranks Frontier LLMs on Biological Research Prompts Open-weights release; 12B active / 120B total parameters; 1M-token context
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation c6013dec-24c1-446d-9c90-83ac80a86860 · outbound
RefusalBench: Why Refusal Rate Misranks Frontier LLMs on Biological Research Prompts SORRY- Bench: Systematically evaluating large lan- guage model safety refusal
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 6b539532-c25e-4445-ad80-74ef7a44f0c4 · outbound
RefusalBench: Why Refusal Rate Misranks Frontier LLMs on Biological Research Prompts SORRY-Bench: Systematically Evaluating Large Language Model Safety Refusal
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation ffa28ec7-3751-42ba-a964-4e66f42c28d6 · outbound
RefusalBench: Why Refusal Rate Misranks Frontier LLMs on Biological Research Prompts Content Analysis: An Introduction to Its Methodology
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 00fff7df-d088-4920-ab28-8219d46a02cc · outbound
RefusalBench: Why Refusal Rate Misranks Frontier LLMs on Biological Research Prompts The art of saying no: Contextual non- compliance in language models
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 1b504a5b-3776-4a91-9246-4843a4b2df69 · outbound
RefusalBench: Why Refusal Rate Misranks Frontier LLMs on Biological Research Prompts The Art of Saying No: Contextual Noncompliance in Language Models
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation b5d7b3ba-568f-4d42-aa10-7fa20ca245b4 · outbound
RefusalBench: Why Refusal Rate Misranks Frontier LLMs on Biological Research Prompts Constitutional AI: Harmlessness from AI Feedback
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation cd4dd58d-1202-454f-a890-7ae279186d4c · outbound
RefusalBench: Why Refusal Rate Misranks Frontier LLMs on Biological Research Prompts Anthropic’s responsible scaling policy
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 6d23fe70-0226-4931-91fe-1975d52042fc · outbound
RefusalBench: Why Refusal Rate Misranks Frontier LLMs on Biological Research Prompts A., Mathur, S., Salabert, D., Ballot, J., R´egulo, C., Metcalfe, T
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation da7d68e0-3579-4211-b94c-d698de212477 · outbound
RefusalBench: Why Refusal Rate Misranks Frontier LLMs on Biological Research Prompts Dual use of ar- tificial intelligence-powered drug discovery
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 4bce7a55-a6b8-43ed-814e-3a9912cdd7a3 · outbound
RefusalBench: Why Refusal Rate Misranks Frontier LLMs on Biological Research Prompts Unresolved cited work
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation a5dcc4fd-ec7f-48e4-9b1a-0b0cac51579a · outbound
RefusalBench: Why Refusal Rate Misranks Frontier LLMs on Biological Research Prompts Training language models to follow instructions with hu- man feedback
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation a9b5081e-678c-4a61-bda6-d4b89781c5c7 · outbound
RefusalBench: Why Refusal Rate Misranks Frontier LLMs on Biological Research Prompts Training language models to follow instructions with human feedback
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation a9a33a66-c8c8-4fb4-be63-55d036d2914b · outbound
RefusalBench: Why Refusal Rate Misranks Frontier LLMs on Biological Research Prompts Harm- Bench: A standardized evaluation framework for automated red teaming and robust re- fusal
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 07c00f47-c693-4eea-96ce-da6e2eed708c · outbound
RefusalBench: Why Refusal Rate Misranks Frontier LLMs on Biological Research Prompts HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust Refusal
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation a356293f-7cbf-43d2-9107-6d4e18597a37 · outbound
RefusalBench: Why Refusal Rate Misranks Frontier LLMs on Biological Research Prompts The WMDP Benchmark: Measuring and Reducing Malicious Use With Unlearning
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 910685a1-f3f7-4069-97bd-72855cabab83 · inbound
BioSecBench-Refusal: A paired metric for performance and alignment in agentic biosecurity risk assessment RefusalBench: Why Refusal Rate Misranks Frontier LLMs on Biological Research Prompts
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9e8cf392-3063-4472-94f6-bd8b27376ef6 · inbound
BioSecBench-Refusal: A paired metric for performance and alignment in agentic biosecurity risk assessment RefusalBench: Why Refusal Rate Misranks Frontier LLMs on Biological Research Prompts
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.