Finite-state tools add only log|M| bits to finite-precision recurrent controllers, while a single tape tool yields Turing completeness with O(log|Q|+log|Γ|) bits, realized exactly by one-layer selective SSMs.
Title resolution pending
12 Pith papers cite this work. Polarity classification is still indexing.
citation-role summary
citation-polarity summary
roles
method 1polarities
use method 1representative citing papers
Embodied CAD deploys solver-grounded LLM agents with stratified L0-L4 skills, action grammar, and solver rewards to achieve high executable rates on multi-step mechanical and mold assembly tasks.
Attention analysis shows that LLM tool selection failures occur at the readout/decision stage, not because the model fails to attend to the correct tool definition.
CoCoDA co-evolves a typed compositional DAG of primitive and composite tools with the agent planner, using signature-based retrieval and a size-based reward to scale libraries efficiently and let an 8B model match or beat a 32B model on math and code benchmarks.
PaperPilot induces editable DAG workflows of paper-search operators and, after workflow imitation plus preference training, lifts a 9B multi-turn agent from 58 to 77 Hit@5 while cutting execution errors to 0%.
NeuroMAS reframes multi-agent language systems as neural architectures where LLM agents learn coordination via reinforcement learning rather than predefined roles.
A single consistency instruction with harmful prior actions causes aligned frontier LLMs to select unsafe options at 91-98% rates in high-stakes domains, with escalation and inverse scaling by model size.
ANNEAL uses Failure-Driven Knowledge Acquisition to localize faults, generate constrained symbolic patches, and validate them before committing to a process knowledge graph, eliminating recurring failures where baselines do not.
HELM raises long-horizon VLA success from 58.4% to 81.5% on LIBERO-LONG by combining episodic memory retrieval, learned failure prediction, and replanning, outperforming context extension or adaptation alone.
Introduces a benchmarking suite with common workload adapters, event schemas, and an evidence gate connecting WebArena Verified, SWE-Gym, and MiniWoB++ for tool-using agents.
SPREG detects logical failures in LLM long-chain reasoning through real-time entropy spikes and performs structured plan repairs using historical distributions, reporting a 20% absolute accuracy gain on AIME25.
citing papers explorer
-
When Does Tool Use Increase the Expressive Power of Finite-Precision Recurrent Models?
Finite-state tools add only log|M| bits to finite-precision recurrent controllers, while a single tape tool yields Turing completeness with O(log|Q|+log|Γ|) bits, realized exactly by one-layer selective SSMs.
-
Embodied CAD: Solver-Grounded LLM Agents for Parametric B-Rep Assembly Modeling
Embodied CAD deploys solver-grounded LLM agents with stratified L0-L4 skills, action grammar, and solver rewards to achieve high executable rates on multi-step mechanical and mold assembly tasks.
-
Looking Is Not Picking: An Attention-Segment Account of Tool-Selection Failures in LLM Agents
Attention analysis shows that LLM tool selection failures occur at the readout/decision stage, not because the model fails to attend to the correct tool definition.
-
CoCoDA: Co-evolving Compositional DAG for Tool-Augmented Agents
CoCoDA co-evolves a typed compositional DAG of primitive and composite tools with the agent planner, using signature-based retrieval and a size-based reward to scale libraries efficiently and let an 8B model match or beat a 32B model on math and code benchmarks.
-
Multi-Turn Agentic Scientific Literature Search via Workflow Induction
PaperPilot induces editable DAG workflows of paper-search operators and, after workflow imitation plus preference training, lifts a 9B multi-turn agent from 58 to 77 Hit@5 while cutting execution errors to 0%.
-
NeuroMAS: Multi-Agent Systems as Neural Networks with Joint Reinforcement Learning
NeuroMAS reframes multi-agent language systems as neural architectures where LLM agents learn coordination via reinforcement learning rather than predefined roles.
-
History Anchors: How Prior Behavior Steers LLM Decisions Toward Unsafe Actions
A single consistency instruction with harmful prior actions causes aligned frontier LLMs to select unsafe options at 91-98% rates in high-stakes domains, with escalation and inverse scaling by model size.
-
ANNEAL: Adapting LLM Agents via Governed Symbolic Patch Learning
ANNEAL uses Failure-Driven Knowledge Acquisition to localize faults, generate constrained symbolic patches, and validate them before committing to a process knowledge graph, eliminating recurring failures where baselines do not.
-
HELM: Harness-Enhanced Long-horizon Memory for Vision-Language-Action Manipulation
HELM raises long-horizon VLA success from 58.4% to 81.5% on LIBERO-LONG by combining episodic memory retrieval, learned failure prediction, and replanning, outperforming context extension or adaptation alone.
-
An Executable Benchmarking Suite for Tool-Using Agents
Introduces a benchmarking suite with common workload adapters, event schemas, and an evidence gate connecting WebArena Verified, SWE-Gym, and MiniWoB++ for tool-using agents.
-
SPREG: Structured Plan Repair with Entropy-Guided Test-Time Intervention for Large Language Model Reasoning
SPREG detects logical failures in LLM long-chain reasoning through real-time entropy spikes and performs structured plan repairs using historical distributions, reporting a 20% absolute accuracy gain on AIME25.
- AppAgent: Multimodal Agents as Smartphone Users