Multi-agent LLM reaches 79% on penetration testing benchmarks
Fine-tuned model drives agents through reconnaissance, scanning, and exploitation without human input on standard tests
Cryptography and Security
Covers all areas of cryptography and security including authentication, public key cryptosytems, proof-carrying code, etc. Roughly includes material in ACM Subject Classes D.4.6 and E.3.
sort pith recommended most recent
Fine-tuned model drives agents through reconnaissance, scanning, and exploitation without human input on standard tests
A lifecycle map shows why provenance and versioning must be built in from the first store.
SEED dataset of 90K sequentially edited images finds high-frequency wavelets help order manipulations even after degradation
· “SEED: A Large-Scale Benchmark for Provenance Tracing in Sequential Deepfake Facial Edits”
Rational attackers profit by accelerating hardware to capture MEV, so delays must exceed cost-based thresholds derived from an optimal-stop
· “Economic Security of VDF-Based Randomness Beacons: Models, Thresholds, and Design Guidelines”
Street-legal attachment reaches 18 percent impersonation rate for under 100 dollars without altering plates
· “Street-Legal Physical-World Adversarial Rim for License Plates”
Query-time and corpus-time filters cut attack success from 67% to 14% while clean QA utility stays put.
Knowing the agent's tools and task before injecting beats payload-only red teaming across models and benchmarks.
· “Know Your Agent: Reconnaissance-Driven Pentesting of AI Agents”
A new taxonomy of why patches fail and why current detection tools score below chance at spotting incomplete fixes.
A local agent chooses redaction, abstraction, replacement, or decoy noise so providers see less and users still get useful replies
RogueMerge jointly optimizes for post-merge success and uses a Taylor approximation to handle unknown weights and new prompts.
· “RogueMerge: Robust and Unified Attacks against LLM Model Merging”
Generator proposes links; symbolic checks on surviving telemetry cut hallucinations to 6 percent while preserving 84 percent precision after
· “HunterAgent: Neuro-Symbolic Attack Trace Reconstruction under Anti-Forensics”
Execution trace analysis lets models create YARA rules that reveal hidden behaviors and classify 47% more families than standard platforms.
· “A Large Language Model Approach to Generating Bypass Rules for Malware Evasion in Analysis Sandbox”
Stateful redirection keeps 64-bit messages recoverable with little change to the generated text
· “Every Bit, Everywhere, All at Once: A Binomial Multibit LLM Watermark”
Conflicting relations merge into the same anchor as benign data, letting attackers steer agent answers on targeted tasks at 93.8 percent avg
· “ShadowMerge: A Novel Poisoning Attack on Graph-Based Agent Memory via Relation-Channel Conflicts”
Fixed rules and deterministic checker keep baseline recall stable across 200 incidents and three reruns in a financial SOC case study.
· “SOCpilot: Verifying Policy Compliance for LLM-Assisted Incident Response”
Attack aligns permuted activations from ~$1 of queries and recovers weights within low L1 error.
· “On the (In-)Security of the Shuffling Defense in the Transformer Secure Inference”
Robust scatter matrices deliver stronger privacy at comparable utility when influential outliers are present.
· “Data anonymization in the presence of outliers via invariant coordinate selection”
Any continuous evasion attempt must degrade retrieval performance, supplying guarantees independent of attacker strategy
· “MEMSAD: Gradient-Coupled Anomaly Detection for Memory Poisoning in Retrieval-Augmented Agents”
SIR-Bench replays real incidents to test if agents discover new evidence or merely echo alerts.
· “SIR-Bench: Evaluating Investigation Depth in Security Incident Response Agents”
Hybrid architecture uses MDI star topology to integrate information-theoretic security into distributed consensus.
Framework adds practitioner checklist and multi-perspective checks to give closer evaluations and concrete fixes.
Dual-stream signals based on semantics detect and trace attacks while preserving text quality.
Adjusting imaginary DFT parts after normalization embeds signals that survive attacks without storage costs or loss of data utility.
Reasoning models beat standard ones roughly 5 to 1 on accounts compromised, at ~$23 per two-hour run.
Transformations mark input origins so models follow only trusted commands while keeping task performance nearly unchanged.
· “Defending Against Indirect Prompt Injection Attacks With Spotlighting”
Projective-geometry adversary families lift the Ω(Ln) baseline to Ω(L n^{1+1/d}) bits for agreement and broadcast.
· “Multivalued Consensus: General Adversaries Require More Communication”
Derandomization hardness plus succinct arguments forces a cheating prover to store nearly every bit.
Protocols on over five billion devices accept complex untrusted data without pairing, exposing reachable denial-of-service and bypass paths.
It accepts the first sufficiently successful mechanism from a sequence, rejects the rest, and uses a bounded number of evaluations while保持纯ε
First study of real-world remote servers finds 325 total issues, dynamic client registration flaws in 96.6 percent.
· “A First Measurement Study on Authentication Security in Real-World Remote MCP Servers”
Six of seven tested accelerators from major vendors allow malicious apps to perform privileged operations.
· “Speed Kills: Exploring Confused Deputy Attacks Through Edge AI Accelerators”
Fixed member and non-member datasets plus two corrections for distribution shift produce valid privacy bounds.
Sequential patching without retraining creates order-dependent conflicts that reverse prior protections in shared layers.
· “Defenses at Odds: Measuring and Explaining Defense Conflicts in Large Language Models”
Attacks using shadows or wet roads suppress 57% of detections and inject false boundaries where pixel methods fail entirely.
· “Systematic Discovery of Semantic Attacks in Online Map Construction through Conditional Diffusion”
ExploitBench measures 16 progressive flags on 41 V8 vulnerabilities and finds public models stop at crashes while private models reach code
· “ExploitBench: A Capability Ladder Benchmark for LLM Cybersecurity Agents”
Staging them as MEV events like sandwiches or arbitrage leaves sender and receiver unlinked under standard tools.
· “Extending Blockchain Untraceability with Plausible Deniability”
Chassis heat patterns identify applications at over 90 percent accuracy from a meter away using 10 seconds of data.
· “ThermalTap: Passive Application Fingerprinting in VR Headsets via Thermal Side Channels”
Natural language goals drive attack selection and composition, unifying tests for ML and generative models with high success rates in case演示
· “Redefining AI Red Teaming in the Agentic Era: From Weeks to Hours”
An evolving attack playbook outperforms fixed jailbreaks and agentic baselines while using fewer tool calls.
· “RedEvoAgent: Automatic Red-Teaming Agent with Experience-Driven Skill Evolution”
A benchmark shows why security-scanner accuracy and availability must be reported separately.
· “Beyond F1: Evaluating Coverage and Failure Recovery in AI Model Security Scanners”
Delegation, goals, and scheduled tasks relabel attacker content as user intent, defeating both the model and permission review.
· “When Context Gets Root: Privilege Escalation in LLM Harnesses”
A reverse-trained backdoor hides its trigger signal but keeps target bias and evades detection.
· “Low-ASR Backdoors: Exploiting Attack Success Rate Reduction and Attacker-Defender Asymmetry”
A round-robin attack across sub-banks cuts mean time to failure from 13 years to about 1 second.
· “From Fleet to Lab: Revisiting the Security and Complexity of Industrial Rowhammer Mitigation”
One epoch of clean-data fine-tuning, no triggers needed, keeps functional correctness and often lifts pass rates.
· “RTLGuard: A Lightweight Teacher-Student Defense for Poisoned RTL Code Generation Models”
An attack built on a small open model transfers to closed-source embedders and LLMs.
· “Vulnerable Code Search: Transferable Attack for Code Language Models”
It stores the attack method, not the topic, and applies that knowledge to future inputs with no parameter updates.
· “A Self-Evolving Multi-Agent Framework Defense against LLM Jailbreak Attacks”
Even specification-valid documents can show a benign value in Office while an extractor feeds a planted trap fact to the model.
· “Beyond the Editing Canvas: Evidence Divergence in OOXML-to-LLM Ingestion”
Visual AI checks 44-pixel tap targets; a proxy scores the URLs those clicks trigger.
· “A Hybrid Security Framework for Mini-Programs: Visual UI Compliance and Network Risk Assessment”
Retraining-free defense for API-only agents: execution attack success drops to 14.5 percent with a per-class skill.
· “SkillShield: Prompt-Space Security Skills for LLM Coding Agents”
The fix already exists: most workload clusters contain a secure reference; copying it lifts scores 60.4%.
One retrieval can seed a worm of agent-written copies that survives removal of the original malicious skill.
Five demonstrations show context manipulation before signing makes every mandate valid but unintended.
· “Beyond the Mandate: A Systematic Security Analysis of the Agent Payments Protocol (AP2)”
A reproducible scenario model separates mechanism-backed RSA risk from conditional PQC risk, with updateable dates for migration planning.
The attack maps model weights onto a chip's real bit-flip sites, reaching near 90 percent success.
· “ROBBIN: Rowhammer-Based Backdoor Injection during Inference”
Zero-oracle SHIELD defense cuts it to 42.7%, but no frontier model is robust to staged server compromise.
· “TrustShiftProbe: Characterizing, Benchmarking, and Defending Staged Trust Attacks on MCP Servers”
Stitch links incoming and outgoing flows via process labels on one host, cutting false positives below 0.01 percent.
· “Effective Pivot Attack Detection via System and Network Information”
Most written rules stay prose for the model to interpret; nothing marks which rules are enforced.
· “When "Do Not" Is Not Deny: Security Rules in CLAUDE.md vs Built-In Controls”
On the SWaT water plant, 1% poisoned training data crushes top clean detectors while PCA and SVM hold.
Rehearsal-free framework hits 61% final-session accuracy and halves forgetting
· “Adapter-Based Few-Shot Continual Learning for Malicious Packet Recognition”
No memory access needed: one injected page steers Qwen2.5 answers 76.6% of the time once retrieved.
· “InjecMEM: Memory Injection Attack on LLM Agent Memory Systems”
On cybersecurity events and two e-commerce churn tasks, rule-aware graphs rank rare anomalies best.
If the validator graph survives any f failures, encrypted sensor data keeps flowing; otherwise it halts.
A spectral trace-of-inverse bound stays nontrivial for deterministic encoders and covers perceptual metrics.
· “Spectrum-Aware Bounds on Invertibility for Privacy-Enhancing Instance Encoding”
Garbled, script-mixed prompts roughly double leaked bits per token, so static jitter thresholds understate exfiltration risk.
· “Adversarial Entropy Inflation Against Gumbel-Based Inference Verification”
A new meta-interface lets trusted administrators install on-device storage policies even when the OS is hostile.
· “SxSSD: A Secure and Extensible Software-defined Solid State Drive”
Eight automated tests map pass/fail verdicts to CRA and NIS2 rules, with under 10 ms added latency.
· “CERTIoT-6G: Continuous Cybersecurity Certification for IoT Devices in 5G/6G Networks”
FIDES reconciles what a model says, codes, and earns: 32 strategies claimed market-beating edge, one delivered it.
· “FIDES: A Concordance Protocol for LLM-Generated Trading Strategies”
Raw HTML, screenshots, TLS, and DNS stay inspectable so researchers can re-derive features as phishing evolves.
· “PhiShark2026: A Multi-Layer Active-Web Raw-Evidence Dataset for Phishing Website Research”
Eleven agents turn passive forum watching into targeted questioning and pull out extra threat intelligence.
· “Towards Automated Cyber Threat Intelligence Elicitation in Underground Forums”
All three models identify high-probability differentials with no false positives in SIMON and SIMECK.
· “Graph Representation Learning of Lightweight IoT Ciphers”
Non-heuristic analysis of Railgun and Hinkal finds thousands of withdrawals with ten or fewer possible source addresses.
· “The Anonymity Gap: Understanding Real Privacy in Shielded UTXO-based Protocols for DeFi”
The paper turns the "no positive-utility attack" test into model checking, with computable thresholds for payments and voting.
A passive probe near the NIC tells benign web use from floods, scans, and probes without packet or host access.
A runtime derives from its execution record what a checkpoint, fork, restore, or merge must preserve, making safety exact.
· “When Can Agents Safely Checkpoint, Fork, Restore, and Merge? Exact Checking for Execution Edits”
A new review maps the crypto value exposed to Shor's algorithm and the post-quantum upgrades already being tested.
· “Cryptocurrencies in the Quantum Age: Migration Paths to PQC”
In constrained Best-of-N sampling, unsafe outputs that pass a learned safety filter can dominate selection when they have heavier reward…
· “Safety Hacking in Constrained Best-of-N Inference-time Scaling”
Synthesizing semantic specs with an LLM and verifying them against benign traffic beats one-shot rule generation.
AgentFlow enforces flow and path policies over labeled data edges in LLM agent systems, reporting 0% confirmed compromise on AgentDojo and…
· “AgentFlow: A Flow-Centric Policy Language and Framework for Securing LLM Agent Systems”
A query-efficient membership inference attack for diffusion models, DIME, derives a test statistic from the optimal denoiser's…
· “DIME: Query-Efficient Framework for Membership Inference on Diffusion Models”
A blockchain token binding decides who may use a device, so access expires and revokes without touching the Bluetooth stack.
· “A Study of Bluetooth Access Control Based on NFT Soft Pairing”