SafeRx-Agent introduces the first fine-grained medication recommendation task at fourth-level ATC codes and a knowledge-grounded multi-agent system that improves prediction accuracy while controlling safety risks on MIMIC datasets.
The aloe family recipe for open and specialized healthcare llms
4 Pith papers cite this work. Polarity classification is still indexing.
citation-role summary
citation-polarity summary
years
2026 4roles
background 1polarities
background 1representative citing papers
NeWTral is a non-linear weight translation framework using MoE routing that reduces average attack success rate from 70% to 13% on unsafe domain adapters across Llama, Mistral, Qwen, and Gemma models up to 72B while retaining 90% knowledge fidelity.
Most of 72 tested LLMs complied with harmful medical-robot orders over half the time, with open-weight models far worse than proprietary ones.
HERALD selectively encrypts sensitive tokens via medical NER, POS policies, and deterministic ciphertext substitution to enable privacy-preserving clinical LLM use while recovering near-plaintext task performance.
citing papers explorer
-
SafeRx-Agent: A Knowledge-Grounded Multi-Agent Framework for Safe and Explainable Medication Recommendation
SafeRx-Agent introduces the first fine-grained medication recommendation task at fourth-level ATC codes and a knowledge-grounded multi-agent system that improves prediction accuracy while controlling safety risks on MIMIC datasets.
-
You Snooze, You Lose: Automatic Safety Alignment Restoration through Neural Weight Translation
NeWTral is a non-linear weight translation framework using MoE routing that reduces average attack success rate from 70% to 13% on unsafe domain adapters across Llama, Mistral, Qwen, and Gemma models up to 72B while retaining 90% knowledge fidelity.
-
Benchmarking the Safety of Large Language Models for Robotic Health Attendant Control
Most of 72 tested LLMs complied with harmful medical-robot orders over half the time, with open-weight models far worse than proprietary ones.
-
Selective Token-Level Cryptographic Redaction for Privacy-Preserving Clinical Deployment of Large Language Models
HERALD selectively encrypts sensitive tokens via medical NER, POS policies, and deterministic ciphertext substitution to enable privacy-preserving clinical LLM use while recovering near-plaintext task performance.