Finite-state tools add only log|M| bits to finite-precision recurrent controllers, while a single tape tool yields Turing completeness with O(log|Q|+log|Γ|) bits, realized exactly by one-layer selective SSMs.
Title resolution pending
12 Pith papers cite this work. Polarity classification is still indexing.
citation-role summary
citation-polarity summary
roles
method 1polarities
use method 1representative citing papers
Embodied CAD deploys solver-grounded LLM agents with stratified L0-L4 skills, action grammar, and solver rewards to achieve high executable rates on multi-step mechanical and mold assembly tasks.
Attention analysis shows that LLM tool selection failures occur at the readout/decision stage, not because the model fails to attend to the correct tool definition.
CoCoDA co-evolves a typed compositional DAG of primitive and composite tools with the agent planner, using signature-based retrieval and a size-based reward to scale libraries efficiently and let an 8B model match or beat a 32B model on math and code benchmarks.
PaperPilot induces editable DAG workflows of paper-search operators and, after workflow imitation plus preference training, lifts a 9B multi-turn agent from 58 to 77 Hit@5 while cutting execution errors to 0%.
NeuroMAS reframes multi-agent language systems as neural architectures where LLM agents learn coordination via reinforcement learning rather than predefined roles.
A single consistency instruction with harmful prior actions causes aligned frontier LLMs to select unsafe options at 91-98% rates in high-stakes domains, with escalation and inverse scaling by model size.
ANNEAL uses Failure-Driven Knowledge Acquisition to localize faults, generate constrained symbolic patches, and validate them before committing to a process knowledge graph, eliminating recurring failures where baselines do not.
HELM raises long-horizon VLA success from 58.4% to 81.5% on LIBERO-LONG by combining episodic memory retrieval, learned failure prediction, and replanning, outperforming context extension or adaptation alone.
Introduces a benchmarking suite with common workload adapters, event schemas, and an evidence gate connecting WebArena Verified, SWE-Gym, and MiniWoB++ for tool-using agents.
SPREG detects logical failures in LLM long-chain reasoning through real-time entropy spikes and performs structured plan repairs using historical distributions, reporting a 20% absolute accuracy gain on AIME25.
citing papers explorer
-
CoCoDA: Co-evolving Compositional DAG for Tool-Augmented Agents
CoCoDA co-evolves a typed compositional DAG of primitive and composite tools with the agent planner, using signature-based retrieval and a size-based reward to scale libraries efficiently and let an 8B model match or beat a 32B model on math and code benchmarks.