Pith. sign in

REVIEW 23 cited by

InternAgent: When Agent Becomes the Scientist -- Building Closed-Loop System from Hypothesis to Verification

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2505.16938 v3 pith:24TDR375 submitted 2025-05-22 cs.AI cs.CLcs.CV

InternAgent: When Agent Becomes the Scientist -- Building Closed-Loop System from Hypothesis to Verification

classification cs.AI cs.CLcs.CV
keywords internagentresearchscientificfieldshoursacrossclosed-loopefficiency
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
Share X Bluesky LinkedIn Reddit HN
read the original abstract

Artificial Intelligence (AI) is accelerating the transformation of scientific research paradigms, not only enhancing research efficiency but also driving innovation. We introduce InternAgent, a unified closed-loop multi-agent framework to conduct Autonomous Scientific Research (ASR) across various scientific research fields, enabling researchers to tackle complicated problems in these fields with unprecedented speed and precision. InternAgent highlights three key advantages: 1) Scalability: InternAgent has demonstrated its versatility across 12 scientific research tasks, capable of generating innovative ideas to enhance the performance of baseline code. 2) Interactivity: InternAgent provides an interface for human expert feedback and multi-agent interaction in automated end-to-end processes, allowing for the seamless integration of domain expert knowledge. 3) Efficiency: InternAgent has achieved promising performance gains in several scientific fields with significantly less time cost compared to human efforts. For instance, in reaction yield prediction, it increased from 27.6% to 35.4% in just 12 hours; in enhancer activity prediction, accuracy rose from 0.65 to 0.79 with only 4 hours of processing; and in 2D semantic segmentation, precision advanced from 78.8% to 81.0% in a mere 30 hours.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 23 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Structure-guided taxonomic placement of divergent RNA viruses with ViraClass

    q-bio.QM 2026-06 unverdicted novelty 7.0

    ViraClass is a hierarchical classifier that uses RdRp protein structure for rank-by-rank taxonomic placement of divergent RNA viruses, outperforming sequence methods at deep evolutionary distances.

  2. Knows: Agent-Native Structured Research Representations

    cs.AI 2026-04 conditional novelty 7.0

    Knows uses a YAML sidecar specification to provide structured, agent-consumable representations of research papers, yielding large accuracy gains for small LLMs on comprehension tasks and rapid community adoption via ...

  3. SeekBrain: An Autonomous Multi-Agent System for Accelerating Neuroscience Discovery

    cs.MA 2026-07 conditional novelty 6.0

    SeekBrain, a multi-agent system with a neuroscience analysis recipe library, outperforms Claude Code and Codex on 32 expert-scored neuroscience analysis tasks and carries out two published-data analyses.

  4. One Reflection Is Not Enough: Self-Correcting Autonomous Research via Multi-Hypothesis Failure Attribution

    cs.AI 2026-06 unverdicted novelty 6.0

    SAGE with MHFA improves failure recovery in autonomous research agents, raising metrics-bearing outputs from 42% to 92% on a 12-topic benchmark versus single-reflection baselines.

  5. Investigating Novice Researchers' Perceptions of Research Privacy Within LLM-Assisted Workflows

    cs.HC 2026-06 unverdicted novelty 6.0

    Interview study of 44 novice researchers finds privacy fears paradoxically accelerate LLM use for faster publication, with misconceptions about idea value and data dilution, and perceived ineffective mitigations.

  6. NanoResearch: Co-Evolving Skills, Memory, and Policy for Personalized Research Automation

    cs.AI 2026-05 unverdicted novelty 6.0

    NanoResearch introduces a tri-level co-evolving framework of skills, memory, and policy to personalize LLM-powered research automation across projects and users.

  7. AIBuildAI: An AI Agent for Automatically Building AI Models

    cs.AI 2026-04 unverdicted novelty 6.0

    AIBuildAI uses a manager agent and three LLM sub-agents to fully automate AI model development and achieves a 63.1% medal rate on MLE-Bench, matching experienced human engineers.

  8. TREX: Automating LLM Fine-tuning via Agent-Driven Tree-based Exploration

    cs.AI 2026-04 unverdicted novelty 6.0

    TREX automates the LLM training lifecycle via collaborative agents and tree-based exploration, delivering consistent performance gains across 10 real-world fine-tuning tasks in FT-Bench.

  9. Can We Predict Before Executing Machine Learning Agents?

    cs.CL 2026-01 unverdicted novelty 6.0

    LLMs primed with verified data reports predict agent solution quality at 61.5% accuracy, powering a Predict-then-Verify agent that converges 6x faster than execution-only baselines.

  10. CodeDistiller: Automatically Generating Code Libraries for Scientific Coding Agents

    cs.AI 2025-11 conditional novelty 6.0

    CodeDistiller distills 250 materials-science GitHub repositories into vetted code libraries that improve the accuracy and scientific soundness of experiments generated by ASD agents.

  11. Structural Enforcement of Statistical Rigor in AI-Driven Discovery: A Functional Architecture

    cs.SE 2025-11 conditional novelty 6.0

    An AI-Scientist guard architecture combining a Haskell monad for online FDR accounting with declarative scaffolding against data leakage; simulation supports it, but the advertised Lean/SPARK verification is absent fr...

  12. Rethinking Scientific Discovery in the Agentic Era

    cs.CL 2026-07 conditional novelty 5.5

    SCION claims an agentic OS with Research Execution Plans and layered memory that beats autonomous research-agent baselines on reading, ideation, molecule design, and antibody screening.

  13. Agon: An Autonomous Large-Scale Omnidisciplinary Research System Built on Prompt Economy

    cs.SE 2026-06 unverdicted novelty 5.0

    Agon is a new autonomous research system using prompt economy loops across 444 iterations to demonstrate scalable omnidisciplinary research and a taxonomy separating machine-fixable failures from those needing human judgment.

  14. AIBuildAI-2: A Knowledge-Enhanced Agent for Automatically Building AI Models

    cs.AI 2026-05 unverdicted novelty 5.0

    AIBuildAI-2 introduces a knowledge-enhanced agent with a hierarchical evolving external knowledge base that dynamically loads relevant AI development expertise, achieving first place on MLE-Bench at 70.7% medal rate.

  15. pAI/MSc: ML Theory Research with Humans on the Loop

    cs.AI 2026-04 unverdicted novelty 5.0

    pAI/MSc is a customizable multi-agent system that reduces human steering by orders of magnitude when turning a hypothesis into a literature-grounded, mathematically established, experimentally supported manuscript dra...

  16. SciDER: Scientific Data-centric End-to-end Researcher

    cs.AI 2026-03 conditional novelty 5.0

    SciDER is a data-centric multi-agent system that automates ideation, raw-data analysis, experiment coding, and critique, with reported leading results on six scientific-agent benchmarks.

  17. SciDER: Scientific Data-centric End-to-end Researcher

    cs.AI 2026-03 unverdicted novelty 5.0

    SciDER introduces a data-centric end-to-end multi-agent system that automates the scientific research lifecycle from raw data analysis to code execution and outperforms general-purpose agents on three benchmarks via s...

  18. TusoAI: Agentic Optimization for Scientific Methods

    cs.AI 2025-09 unverdicted novelty 5.0

    TusoAI is an LLM-based agent that builds and iteratively optimizes domain-specific computational methods for scientific data analysis, outperforming expert baselines on RNA-seq denoising and earth monitoring while rep...

  19. MLEvolve: A Self-Evolving Framework for Automated Machine Learning Algorithm Discovery

    cs.AI 2026-06 unverdicted novelty 4.0

    MLEvolve is a self-evolving multi-agent LLM system with Progressive MCGS, Retrospective Memory, and adaptive coding modes that reports SOTA medal and submission rates on MLE-Bench under a 12-hour budget while outperfo...

  20. SciAtlas: A Large-Scale Knowledge Graph for Automated Scientific Research

    cs.AI 2026-05 unverdicted novelty 4.0

    SciAtlas builds a large-scale multi-disciplinary academic knowledge graph and a neuro-symbolic retrieval system to support automated scientific research tasks such as literature review and idea positioning.

  21. AI for Auto-Research: Roadmap & User Guide

    cs.AI 2026-05 conditional novelty 4.0

    AI can generate research artifacts faster than it can verify them, so across all eight lifecycle stages the credible deployment mode is human-governed collaboration rather than full autonomy.

  22. AI for Auto-Research: Roadmap & User Guide

    cs.AI 2026-05 unverdicted novelty 4.0

    The paper delivers a stage-by-stage roadmap for AI in research, showing reliable assistance in retrieval and tool tasks but fragility in novelty and judgment, advocating human-governed collaboration.

  23. Evolving Roles of LLMs in Scientific Innovation: Assistant, Collaborator, Scientist, and Evaluator

    cs.DL 2025-07 unverdicted novelty 4.0

    The paper proposes a four-role framework for LLMs in scientific innovation and reviews methods, benchmarks, and limitations across Assistant, Collaborator, Scientist, and Evaluator roles.