PrincipalBench exposes a sharp split in frontier LLMs between selective and over-refusing behavior on multi-party loyalty, with prompt scaffolding and KL distillation reducing harm rates but only along an existing leak/over-refusal trade-off.
Shao Y, Li T, Shi W, Liu Y, Yang D
6 Pith papers cite this work, alongside 4 external citations. Polarity classification is still indexing.
citation-role summary
citation-polarity summary
roles
other 1polarities
unclear 1representative citing papers
MAC-Bench is a new adversarial benchmark that converts legal texts into executable scenarios via the SERV pipeline to measure procedural compliance in multi-agent LLM systems using CSR and MG metrics.
CalBench is a new benchmark for multi-agent LLM calendar scheduling that measures task success, excess cost, communication efficiency, burden fairness, and privacy leakage under private information constraints.
Agent-SafetyBench shows no tested LLM agent exceeds 60% safety score, attributing failures to lack of robustness and risk awareness.
A data-centric survey finds that only information-flow control covers compositional and cross-session leakage in LLM agents and that no single benchmark tests an agent across all its data surfaces under one policy.
A fine-tuned local Qwen 3.5 27B model achieves 95% category-level accuracy on security document classification, outperforming commercial models on both internal and external test sets while keeping processing local.
citing papers explorer
-
Whose Side Is Your Agent On? Multi-Party Principal Loyalty in LLM Agents
PrincipalBench exposes a sharp split in frontier LLMs between selective and over-refusing behavior on multi-party loyalty, with prompt scaffolding and KL distillation reducing harm rates but only along an existing leak/over-refusal trade-off.
-
Beyond Goodhart's Law: A Dynamic Benchmark for Evaluating Compliance in Multi-Agent Systems
MAC-Bench is a new adversarial benchmark that converts legal texts into executable scenarios via the SERV pipeline to measure procedural compliance in multi-agent LLM systems using CSR and MG metrics.
-
CalBench: Evaluating Coordination-Privacy Trade-offs in Multi-Agent LLMs
CalBench is a new benchmark for multi-agent LLM calendar scheduling that measures task success, excess cost, communication efficiency, burden fairness, and privacy leakage under private information constraints.
-
Agent-SafetyBench: Evaluating the Safety of LLM Agents
Agent-SafetyBench shows no tested LLM agent exceeds 60% safety score, attributing failures to lack of robustness and risk awareness.
-
Agents That Know Too Much: A Data-Centric Survey of Privacy in LLM Agents
A data-centric survey finds that only information-flow control covers compositional and cross-session leakage in LLM agents and that no single benchmark tests an agent across all its data surfaces under one policy.
-
Security Document Classification with a Fine-Tuned Local Large Language Model: Benchmark Data and an Open-Source System
A fine-tuned local Qwen 3.5 27B model achieves 95% category-level accuracy on security document classification, outperforming commercial models on both internal and external test sets while keeping processing local.