Factored trees let agentic teams reuse past solution paths across problems
GRAFT-ATHENA embeds decision sequences as metric fingerprints so new physics tasks draw on accumulated experience and invent new solvers.
Multiagent Systems
Covers multiagent systems, distributed artificial intelligence, intelligent agents, coordinated interactions. and practical applications. Roughly covers ACM Subject Class I.2.11.
sort pith recommended most recent
GRAFT-ATHENA embeds decision sequences as metric fingerprints so new physics tasks draw on accumulated experience and invent new solvers.
Analysis of CPU orchestration in autonomous agents yields two methods that raise hardware overlap and balance mixed requests on hybrid CPU-G
· “Towards Understanding, Analyzing, and Optimizing Agentic AI Execution: A CPU-Centric Perspective”
A multi-agent system with global context resolves conflicts and fragmentation to improve retrieval accuracy on complex tasks.
· “MemGraphRAG: Memory-based Multi-Agent System for Graph Retrieval-Augmented Generation”
By comparing memory operations from identical states, LoGo-GRPO supplies precise signals while retaining end-to-end trajectory rewards.
· “Memory-R2: Fair Credit Assignment for Long-Horizon Memory-Augmented LLM Agents”
Production run shows 70% labor reduction and major drops in emissions and water use for document workflows.
· “MADP: A Multi-Agent Pipeline for Sustainable Document Processing with Human-in-the-Loop”
Sampling intermediate states from offline demonstrations lets regularized gradients find lower-exploitability equilibria under fixed compute
· “Data-Augmented Game Starts for Accelerating Self-Play Exploration in Imperfect Information Games”
The benchmark's 145 tasks show the gap is trade-off reasoning and geometry planning, not rule recall.
Inserting an unmodified model at three points in CPFA lets robots collect more resources with less variance when team or arena size changes.
· “LLM-Foraging: Large Language Models for Decentralized Swarm Robot Foraging”
An interactive visual system let participants create slide decks with higher expert-rated coverage and structure while lowering cognitive负荷.
· “MindTrellis: Co-Creating Knowledge Structures with AI through Interactive Visual Exploration”
Iterative aggregation and refinement let the system assemble answers from distributed specialist agents without broadcasting every query.
· “Talk to Right Specialists: Iterative Routing in Multi-agent Systems for Question Answering”
Seller inference recovers budgets nearly one-for-one from natural-language profiles, and confidentiality instructions do not stop it.
· “When Agents Shop for You: Role Coherence in AI-Mediated Markets”
By routing tasks to specialized agents and verifying against physics models, the system generates auditable plans that simulations indicate
MoRe composes learned role vectors per query, giving one model adaptive multi-perspective reasoning in a single pass.
· “One Model, Many Minds: Unlocking Multi-Agent Synergy in a Single Agent via Mixture of Roles”
Across 75 tasks it scores 60 medals, 49 gold, versus 55 and 34 for Claude Code, with auditable lineages.
· “Praxist: From Experimental Artifacts to Solution Lineages”
An LLM with memory and tool access sets laser, temperature, and overlap from past builds and ASTM test results.
· “AI Agentic Selective Laser Sintering Process Optimization”
HypoForge learns from adversarial critique and empirical outcomes, improving both without fine-tuning.
A prompt-only relay with public-test checks at every step closes most of the gap to heavier systems at one-third the cost.
· “MARS: Multi-Specialist LLM Relay System for Competitive Programming”
Bidding, reputation, and subcontracting outperform every baseline across math, code, quizzes, and agent tasks.
· “Markets, Not Planners: Decentralized Orchestration of LLM Agents with Private Information”
The discoveries include a 604-point kissing configuration, a new Kakeya family, and an improved Erdős bound.
· “Autonomous Mathematical Discovery in an Open-World Multi-Agent Environment”
New tests show agents converge after reading each other's solutions; independent generation preserves the gain.
· “The Interaction Tax: When Communication Erases Diversity in Multi-Agent Teams”
Two- and three-worker splits hit 0.830 on VAT cases vs 0.720 and 0.770 at the extremes; one bad record degrades every setup.
· “Right-Sizing LLM-Agent Decomposition in VAT Determination: A Pilot Controlled Sweep”
Ten-week production run also cut fault incidents by more than 60 percent, the paper reports.
Patches prompts, tools, and control logic from failure traces, beating hand-tuned harnesses with far fewer rollouts.
· “AutoSaddler: Automatic Harness Optimization with Durable Updates from Agent Execution Traces”
Gossip's failure population equals message reach times lifetime; pre-filed predictions held on third-party code.
· “Predicting the scale limits of social mechanisms in agent societies”
Run locally on macOS, Windows, and Linux, logging every perception and decision for study.
· “Minimal Local Simulation Foundations for LLM- and VLM-Driven Agents in 2D and 3D Environments”
Informed LLM planners still dispatch 23–29% of faulty steps; a deterministic orchestrator blocks all of them.
· “Physical Agentic AI: An Architecture for Orchestrating a Robot Crew with LLMs”
Correcting a small rigid peak drift more than doubles retrieval; recalibration on measured spectra restores coverage.
Two identical agents settle on one admissible joint behavior from their interaction alone.
· “Opinion-Guided Layered Strategies for Decentralized Coordination”
In crystallization tests, simulator experiments cut vague wording fourfold and lifted expert ratings above 4.
· “LLM Agents Perform Controlled Experiments Using Simulation Models”
Keeping live state in one VM lineage and discarding branch VMs cuts per-task cost by 34–70% on 200 tasks.
TessIndex anchors compact proofs on-chain and keeps rich metadata off-chain, so agents can be verified before payment flows.
· “TessIndex: Capability Verified Identity System for the Agent Economy”
A single task-agnostic optimizer treats agent trajectories as loss signals and improves every backbone it runs on.
Proving that building a full ranking need not cost extra metric distortion over picking a single winner.
One shared-weight supernet plus a cross-hospital prior finds better transformer designs with less compute.
In hospital tests, the forecast-aware rollout serves all requests and shortens the longest waits.
New games make each reasoning depth yield a distinct action, exposing where models fail.
· “Level-k Distinguishable Mechanisms for Evaluating Bounded Rationality in LLMs”
Adversarial tests show shutdown resistance is an incentive problem, so safer agents need better goals, tools, and supervision.
Catching 99.91 percent of flight-critical faults meets the one-in-ten-million per hour crash target.
· “A Safety-Driven Architectural Framework for Fail-Operational Drone Swarms in Critical Missions”
One task can hit a shared LLM backend as a calm sequence, a burst, or a mesh—reasoning gaps fit a log-normal curve.
· “Towards Traffic Modelling of Multi-Agent Systems: The Role of Coordination Topology”
An orchestrator fuses RL, PID, and common-sense rules while the LLM works offline to refine rewards.
Accuracy matches ARG-Designer on six reasoning and code benchmarks while token bills shrink by about a fifth.
When each module reads only its own slot, held-out composition succeeds; fully visible twins mostly memorize.
In 448 trials, ranked peer posts raise wording similarity; four sources show no reliable stance edge.
A controlled test shows the LLM layer gains a small but significant edge only when a mid-run surge creates headroom
A proposed architecture splits each safety-message verdict among the car, roadside edge, and cloud, biasing ambiguity toward escalation.
· “Autonomous Cyber Defense in Connected Vehicles: A Multi-Agent Approach to V2X Security”
Role-specific memory and reflection lift medical exam accuracy several points above GPT-4 and prior multi-agent baselines.
· “Adaptive Memory and Reflection Multi-Agent System for Medical Question Answering”
Five specialist agents feed a shared blackboard so every diagnosis traces back to specific findings and measurements.
· “DentAgent: Evidence-Centric Multi-Agent Coordination for Multimodal Dental Reasoning”
A state-carrying LLM tuner beats five baselines everywhere; removing its memory cuts frontier quality by 58.5%.
· “StateTune: Transforming LLM-Assisted EDA Flow Tuning into a Stateful, Closed-Loop Process”
Automation rankings diverge from assistant rankings; leaderboards miss a separate capability.
· “CentaurBench: Benchmarking LLM Capabilities on Augmenting vs. Automating Real-World Work Tasks”
Training only the recovery steps after a failure point lets a small model overtake its teacher and cut tool calls.
Bayesian partner tracking turns belief into action, closing the belief-action gap from 0.41 to 0.20.
· “Bayesian Partner Modelling enables Adaptive Replanning for LLM Coordination”
Locally deployed agents coordinate plume analysis and reporting with 92% routing, 85% emission-estimate success.
A two-layer scheme lets a central UTM ban unsafe airspace events while each drone optimizes only inside the safe set.
· “Model Predictive Supervisory Control for Hierarchical and Distributed UAS Traffic Management”
A threshold reward share triggers the switch from cheap to capable models; bandit learners match theory closely.
· “Contracting for LLM Delegation: Moral Hazard in Technology and Effort Choice”
A formal model shows iterated cross-agent steps can reach goals no member plans alone, yet goals hidden to all can never be certified.
A new bound shows spatial grouping and time horizon trade off evenly, cutting plan cost by 25x.
· “A Theoretical Framework for Parallel Lifelong MAPF Using Group Decentralized Planning”
Independent review catches bad experiments while generated ideas still improve image-retrieval recall by 1.85 points.
True-state MAE drops from 0.41 to 0.10 as physics supervision improves offline CAV control.
In overloaded Wi-Fi tests, uncompressed DMPC received 0.00% of messages; the LSTM code received 98% and converged.
· “Communication Reduction via Semantic-Based Encoding in DMPC Using LSTMs”
Plan–perceive–act loop clarifies vague requests and anchors edits to 3D geometry for consistent novel views.
· “DesignAgent3D: Interactive 3D Scene Editing via Designer-like Multimodal Reasoning”
A proposed memory controller moves knowledge through active, dormant, and retired states, enabling governance and later reactivation.
· “Towards Reversible Forgetting: Managing Obsolete Knowledge in Continual Enterprise AI Agents”
On a paged-attention task, two agents sharing conclusions reach a 290.8x speedup versus 142.6x for one agent.
· “KernelArc: A Multi-Agent Framework for GPU Kernel Optimization”
Fit to one-step opinion flips, the rule beats baselines and generalizes to network shapes it never saw.
· “Physics of Agents: Statistical Mechanics Predicts Collective Behavior of AI Agents”
A 12,500-run sweep finds no divergent susceptibility; subcritical cascades explain the sharp segregation transition.
· “Absence of critical scaling in the Schelling segregation model”
Distilling public skill diffs into reusable experience lifts all five benchmarks and improves cross-model transfer.
· “VCE-Skill: Enhancing Skill Self-Evolution with Version-Change Experience”
A hypothesis-driven agent using per-instance feedback outranks 96 models and writes a faster motif algorithm.
· “The Little Scientist: LLM Agent-Driven Discovery via the Scientific Method”
A runtime ethics watchdog withholds more answers when imaging or labs are incomplete, and accuracy on answered cases rises.
· “ETHOS: Towards a Modular Ethics Framework for Clinical Multi-Agent Systems”
Each model scores its own confidence, familiarity, and conflict; hard queries get routed to a deeper pipeline.
Mission success rises from 35% to 62% as test-time communication denial increases.
In a 300-firm agent simulation, pooling AI's tail losses keeps firms solvent and narrows adoption gaps.
· “Insurance as AI Risk Infrastructure: A Generative-Agent Simulation of AI Adoption”
Across ten apps and three production traces, non-LLM work dominates latency, memory, and cost; coordinated serving cuts latency 29-40%.
· “From LLM Inference to Agentic Workloads: Characterization and Implications for Serving Systems”
A range platform, an attack agent, and a detector turn exercise outcomes into mutual training signal.
· “SysEvolve: An AI-native, safe, autonomous adversarial attack-defense co-evolutionary system”
Six of nine strategies are weakly dominated; the survivors are all conditional cooperators.
· “The Open-Strategy Dictator Game: Cooperation Under Mutual Transparency”
Policy gradients match categorical execution; KL-mirror updates survive agents joining and leaving.
· “Submodular Policy Learning for Distributed Task Allocation in Open Multi-Agent Systems”
One 2B model trained on five recovery roles restores most task goals lost to drift in AppWorld.
SHAP-guided reward relabeling lets slices trained on random data coordinate without messages.
· “XAI-Guided Conservative Decentralized Execution for Offline Multi-Agent Network Slicing”
Post-training on six negotiation and scheduling games lifts a 4-billion-parameter model to 0.627 average utility, on par with GPT-4.1 and…
Under them, joint communication-control design reduces to Riccati recursions plus a finite schedule search.
MEC-hosted latent translators cut belief error 68 percent at equal communication cost in a 6G case study.
· “Heterogeneity-Aware Belief Synchronization for Semantic Communication in AI-Native 6G Networks”
It formalizes ideas, compares them pairwise against past papers, and Elo-aggregates scores to match expert judgment.
· “LigBench: A Unified and Human-Aligned Benchmark for LLM-based Research Idea Generation”