REVIEW 7 cited by
Control Risk for Potential Misuse of Artificial Intelligence in Science
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
The expanding application of Artificial Intelligence (AI) in scientific fields presents unprecedented opportunities for discovery and innovation. However, this growth is not without risks. AI models in science, if misused, can amplify risks like creation of harmful substances, or circumvention of established regulations. In this study, we aim to raise awareness of the dangers of AI misuse in science, and call for responsible AI development and use in this domain. We first itemize the risks posed by AI in scientific contexts, then demonstrate the risks by highlighting real-world examples of misuse in chemical science. These instances underscore the need for effective risk management strategies. In response, we propose a system called SciGuard to control misuse risks for AI models in science. We also propose a red-teaming benchmark SciMT-Safety to assess the safety of different systems. Our proposed SciGuard shows the least harmful impact in the assessment without compromising performance in benign tests. Finally, we highlight the need for a multidisciplinary and collaborative effort to ensure the safe and ethical use of AI models in science. We hope that our study can spark productive discussions on using AI ethically in science among researchers, practitioners, policymakers, and the public, to maximize benefits and minimize the risks of misuse.
Forward citations
Cited by 7 Pith papers
-
BadScientist: Can a Research Agent Write Convincing but Unsound Papers that Fool LLM Reviewers?
An LLM agent generating fabricated papers without experiments gets acceptance-level scores from LLM reviewers up to 82% of the time, and simple integrity-checking mitigations barely beat random.
-
SciToolAgent: A Knowledge Graph-Driven Scientific Agent for Multi-Tool Integration
SciToolAgent uses a knowledge graph of over 500 scientific tools to help LLMs select and chain tools, reaching 94% accuracy on a new 531-question benchmark.
-
Investigating the Vulnerability of LLM-as-a-Judge Architectures to Prompt-Injection Attacks
Greedy Coordinate Gradient-optimized suffixes appended to one candidate answer flip Qwen2.5-3B and Falcon3-3B judge verdicts in over 30% of MT-Bench pairwise comparisons.
-
Forbidden Science: Dual-Use AI Challenge Benchmark and Scientific Refusal Tests
A new 512-prompt benchmark claims to measure LLM over-refusal on scientific dual-use questions, but its design and labeling flaws undermine the claim.
-
Disrupt Your Research Using Generative AI Powered ScienceSage
A generative-AI research assistant connects document chat, multimodal chat, and report generation through shared knowledge bases, with an evaluation claiming hybrid vector and knowledge-graph retrieval outperforms eit...
-
Challenges in Guardrailing Large Language Models for Science
A position paper proposing a guardrail framework with four dimensions (trustworthiness, ethics & bias, safety, legal) and implementation strategies for scientific LLM use.
-
Alignment and Safety in Large Language Models: Safety Mechanisms, Training Paradigms, and Emerging Challenges
A broad survey of LLM alignment that catalogs objectives, benchmarks, SFT/RLHF/DPO methods, and safety challenges, without contributing new experimental or theoretical results.
Discussion (0). Continue with ORCID to comment.