Pith. sign in

REVIEW 7 cited by

Control Risk for Potential Misuse of Artificial Intelligence in Science

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2312.06632 v1 pith:WW7ZBHOJ submitted 2023-12-11 cs.AI

classification cs.AI
keywords sciencerisksmisusemodelsartificialcontrolharmfulintelligence
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

The expanding application of Artificial Intelligence (AI) in scientific fields presents unprecedented opportunities for discovery and innovation. However, this growth is not without risks. AI models in science, if misused, can amplify risks like creation of harmful substances, or circumvention of established regulations. In this study, we aim to raise awareness of the dangers of AI misuse in science, and call for responsible AI development and use in this domain. We first itemize the risks posed by AI in scientific contexts, then demonstrate the risks by highlighting real-world examples of misuse in chemical science. These instances underscore the need for effective risk management strategies. In response, we propose a system called SciGuard to control misuse risks for AI models in science. We also propose a red-teaming benchmark SciMT-Safety to assess the safety of different systems. Our proposed SciGuard shows the least harmful impact in the assessment without compromising performance in benign tests. Finally, we highlight the need for a multidisciplinary and collaborative effort to ensure the safe and ethical use of AI models in science. We hope that our study can spark productive discussions on using AI ethically in science among researchers, practitioners, policymakers, and the public, to maximize benefits and minimize the risks of misuse.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 7 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. BadScientist: Can a Research Agent Write Convincing but Unsound Papers that Fool LLM Reviewers?

    cs.CR 2025-10 conditional novelty 6.0 of 10

    An LLM agent generating fabricated papers without experiments gets acceptance-level scores from LLM reviewers up to 82% of the time, and simple integrity-checking mitigations barely beat random.

  2. SciToolAgent: A Knowledge Graph-Driven Scientific Agent for Multi-Tool Integration

    cs.AI 2025-07 conditional novelty 4.0 of 10

    SciToolAgent uses a knowledge graph of over 500 scientific tools to help LLMs select and chain tools, reaching 94% accuracy on a new 531-question benchmark.

  3. Investigating the Vulnerability of LLM-as-a-Judge Architectures to Prompt-Injection Attacks

    cs.CL 2025-05 conditional novelty 4.0 of 10

    Greedy Coordinate Gradient-optimized suffixes appended to one candidate answer flip Qwen2.5-3B and Falcon3-3B judge verdicts in over 30% of MT-Bench pairwise comparisons.

  4. Forbidden Science: Dual-Use AI Challenge Benchmark and Scientific Refusal Tests

    cs.CL 2025-02 reject novelty 4.0 of 10

    A new 512-prompt benchmark claims to measure LLM over-refusal on scientific dual-use questions, but its design and labeling flaws undermine the claim.

  5. Disrupt Your Research Using Generative AI Powered ScienceSage

    cs.IR 2025-02 conditional novelty 3.0 of 10

    A generative-AI research assistant connects document chat, multimodal chat, and report generation through shared knowledge bases, with an evaluation claiming hybrid vector and knowledge-graph retrieval outperforms eit...

  6. Challenges in Guardrailing Large Language Models for Science

    cs.AI 2024-11 conditional novelty 3.0 of 10

    A position paper proposing a guardrail framework with four dimensions (trustworthiness, ethics & bias, safety, legal) and implementation strategies for scientific LLM use.

  7. Alignment and Safety in Large Language Models: Safety Mechanisms, Training Paradigms, and Emerging Challenges

    cs.AI 2025-07 reject novelty 1.0 of 10

    A broad survey of LLM alignment that catalogs objectives, benchmarks, SFT/RLHF/DPO methods, and safety challenges, without contributing new experimental or theoretical results.

Pith tools