Introduces DataClawBench benchmark for exploratory financial data analysis by agents and reports that exploration does not reliably improve task outcomes in noisy cross-domain settings.
From mind to machine: The rise of manus ai as a fully autonomous digital agent.arXiv preprint arXiv:2505.02024, 2025
11 Pith papers cite this work. Polarity classification is still indexing.
citation-role summary
citation-polarity summary
years
2026 11roles
background 3polarities
background 3representative citing papers
GTA-2 benchmark shows frontier models achieve below 50% on atomic tool tasks and only 14.39% success on realistic long-horizon workflows, with execution harnesses like Manus providing substantial gains.
QuadAgent uses an asynchronous multi-agent architecture with an Impression Graph for scene memory and vision-based avoidance to enable training-free vision-language guided agile quadrotor flight, outperforming baselines in simulations and achieving real-world speeds up to 5 m/s.
LEADS is an LLM-agent framework that discovers hybrid models for cardiac EP digital twins by treating domain knowledge as an action space, outperforming human-designed and other LLM-based hybrids on synthetic and real data.
HarnessX composes typed harness components, evolves them from execution traces via a four-stage meta-agent pipeline, and jointly fine-tunes the agent model, reporting +14.5% average peak gains on five benchmarks (validated only on the evolution set).
AgenticQwen small models trained via reasoning and agentic RL with dual data flywheels achieve strong benchmark performance and close the gap to larger models on industrial search and data analysis tasks.
GenericAgent outperforms other LLM agents on long-horizon tasks by maximizing context information density with fewer tokens via minimal tools, on-demand memory, trajectory-to-SOP evolution, and compression.
BiasIG is a multi-dimensional benchmark for social biases in T2I models that shows debiasing interventions frequently cause confounding discrimination effects.
A bidirectional traceability tree and multiagent LLM framework completes fragmented IoT rules, raising completion rates by 43% and cutting logical conflicts by over 21%.
Survey framing LLM agents as model-plus-harness systems, decomposing harness responsibilities, mapping them to tasks, and highlighting open challenges in evaluation, safety, and co-evolution.
Self-sovereign agents are AI systems that economically sustain and extend their operation without human involvement; the paper analyzes remaining technical barriers and discusses associated security, societal, and governance challenges.
citing papers explorer
-
DataClawBench: An Agent Benchmark for Exploratory Real-World Financial Data Analysis
Introduces DataClawBench benchmark for exploratory financial data analysis by agents and reports that exploration does not reliably improve task outcomes in noisy cross-domain settings.
-
GTA-2: Benchmarking General Tool Agents from Atomic Tool-Use to Open-Ended Workflows
GTA-2 benchmark shows frontier models achieve below 50% on atomic tool tasks and only 14.39% success on realistic long-horizon workflows, with execution harnesses like Manus providing substantial gains.
-
QuadAgent: A Responsive Agent System for Vision-Language Guided Quadrotor Agile Flight
QuadAgent uses an asynchronous multi-agent architecture with an Impression Graph for scene memory and vision-based avoidance to enable training-free vision-language guided agile quadrotor flight, outperforming baselines in simulations and achieving real-world speeds up to 5 m/s.
-
Learning Cardiac Electrophysiology Digital Twins Through Agentic Discovery of Hybrid Structure
LEADS is an LLM-agent framework that discovers hybrid models for cardiac EP digital twins by treating domain knowledge as an action space, outperforming human-designed and other LLM-based hybrids on synthetic and real data.
-
HarnessX: A Composable, Adaptive, and Evolvable Agent Harness Foundry
HarnessX composes typed harness components, evolves them from execution traces via a four-stage meta-agent pipeline, and jointly fine-tunes the agent model, reporting +14.5% average peak gains on five benchmarks (validated only on the evolution set).
-
AgenticQwen: Training Small Agentic Language Models with Dual Data Flywheels for Industrial-Scale Tool Use
AgenticQwen small models trained via reasoning and agentic RL with dual data flywheels achieve strong benchmark performance and close the gap to larger models on industrial search and data analysis tasks.
-
GenericAgent: A Token-Efficient Self-Evolving LLM Agent via Contextual Information Density Maximization (V1.0)
GenericAgent outperforms other LLM agents on long-horizon tasks by maximizing context information density with fewer tokens via minimal tools, on-demand memory, trajectory-to-SOP evolution, and compression.
-
BiasIG: Benchmarking Multi-dimensional Social Biases in Text-to-Image Models
BiasIG is a multi-dimensional benchmark for social biases in T2I models that shows debiasing interventions frequently cause confounding discrimination effects.
-
Exploring and Complementing End Users' Requirements in IoT enabled System
A bidirectional traceability tree and multiagent LLM framework completes fragmented IoT rules, raising completion rates by 43% and cutting logical conflicts by over 21%.
-
From Question Answering to Task Completion: A Survey on Agent System and Harness Design
Survey framing LLM agents as model-plus-harness systems, decomposing harness responsibilities, mapping them to tasks, and highlighting open challenges in evaluation, safety, and co-evolution.
-
Self-Sovereign Agent
Self-sovereign agents are AI systems that economically sustain and extend their operation without human involvement; the paper analyzes remaining technical barriers and discusses associated security, societal, and governance challenges.