EvoVuln evolves executable detection policies for five smart-contract vulnerability types using cold-start synthetic testing followed by few-shot refinement on five vulnerable and five safe contracts, reaching 71% macro F1 and enabling a small model to beat a large zero-shot model by 19 points at un
Toolformer: Language models can teach themselves to use tools
8 Pith papers cite this work. Polarity classification is still indexing.
citation-role summary
citation-polarity summary
years
2026 8roles
background 1polarities
background 1representative citing papers
AgentTether repairs 69% of initially failed LLM agent tasks on τ-bench by combining graph-guided root-cause diagnosis, cross-iteration repair memory, and guarded runtime intervention, improving over blind retry by 26 percentage points.
Multi-agent LLM frameworks can spread compromises across agent boundaries via insecure memory inheritance during subagent spawning.
Case study of CMBAgent on 18 astrophysical tasks finds strong performance on well-specified problems but frequent silent failures yielding physically inconsistent outputs.
AgentGate decomposes routing into action decision and structural grounding stages, allowing small 3B-7B models to dispatch queries competitively on a curated benchmark after targeted fine-tuning.
A QoS-aware constrained packer restores 100% skill deliverability under token/risk/tool budgets at ~1.14 hit-rate points cost over unconstrained Top-5 on 35k skills.
CyberAId is a proposed on-premise multi-agent system that coordinates LLM subagents with classical security tools to improve threat response and regulatory alignment in financial services.
A per-customer multi-head attention transformer predicts data center SLA breaches 30 minutes ahead by encoding rules as JSON for training and emitting role-specific outputs for finance, operations, and compliance.
citing papers explorer
-
Knowledge Over Parameters: Evolving Smart Contract Vulnerability Detection
EvoVuln evolves executable detection policies for five smart-contract vulnerability types using cold-start synthetic testing followed by few-shot refinement on five vulnerable and five safe contracts, reaching 71% macro F1 and enabling a small model to beat a large zero-shot model by 19 points at un
-
AgentTether: Graph-Guided Diagnosis and Runtime Intervention for Reliable LLM Agent Operation
AgentTether repairs 69% of initially failed LLM agent tasks on τ-bench by combining graph-guided root-cause diagnosis, cross-iteration repair memory, and guarded runtime intervention, improving over blind retry by 26 percentage points.
-
When Child Inherits: Modeling and Exploiting Subagent Spawn in Multi-Agent Networks
Multi-agent LLM frameworks can spread compromises across agent boundaries via insecure memory inheritance during subagent spawning.
-
Plausible but Wrong: A case study on Agentic Failures in Astrophysical Workflows
Case study of CMBAgent on 18 astrophysical tasks finds strong performance on well-specified problems but frequent silent failures yielding physically inconsistent outputs.
-
AgentGate: A Lightweight Structured Routing Engine for the Internet of Agents
AgentGate decomposes routing into action decision and structural grounding stages, allowing small 3B-7B models to dispatch queries competitively on a curated benchmark after targeted fine-tuning.
-
SkillSelect-Serve: QoS-Aware Budgeted Skill Service Recommendation for LLM Agents
A QoS-aware constrained packer restores 100% skill deliverability under token/risk/tool budgets at ~1.14 hit-rate points cost over unconstrained Top-5 on 35k skills.
-
CyberAId: AI-Driven Cybersecurity for Financial Service Providers
CyberAId is a proposed on-premise multi-agent system that coordinates LLM subagents with classical security tools to improve threat response and regulatory alignment in financial services.
-
A Multi-Head Attention Approach for SLA Compliance Monitoring in Data Centers
A per-customer multi-head attention transformer predicts data center SLA breaches 30 minutes ahead by encoding rules as JSON for training and emitting role-specific outputs for finance, operations, and compliance.