SCICONVBENCH is a new benchmark evaluating LLMs on multi-turn disambiguation and inconsistency resolution for task formulation in computational science, with frontier models reaching only 52.7% success on fluid mechanics disambiguation cases.
Honeycomb: A flexible llm-based agent system for materials science
9 Pith papers cite this work, alongside 5 external citations. Polarity classification is still indexing.
citation-role summary
citation-polarity summary
representative citing papers
TS-Reasoner is a domain-oriented agent using LLMs, computational tools, and error feedback for multi-step time series inference, showing better performance than general LLMs on understanding and reasoning benchmarks.
ALL-FEM fine-tunes LLMs on a corpus of verified FEniCS scripts and uses multi-agent workflows to automate finite element code generation, achieving 71.79% success on 39 benchmarks across elasticity, flow, and coupled problems.
Evo-Memory is a new streaming benchmark and evaluation framework for self-evolving memory in LLM agents, unifying over ten memory modules and introducing the ReMem pipeline for continual improvement on multi-turn and reasoning datasets.
IR-Agent is a multi-agent LLM framework that emulates expert IR spectral analysis procedures to improve molecular structure elucidation accuracy and adaptability.
SciHorizon-DataEVA is a multi-agent system that applies Sci-TQA2 principles across four dimensions to assess AI-readiness of heterogeneous scientific data via dynamic profiling and self-correcting evaluation workflows.
A fine-tuned LLM called Perovskite-R1, built from curated perovskite literature and material libraries, proposes precursor additives and designs with some experimental validation showing improved stability and performance.
A survey consolidating benchmarks, agent frameworks, real-world applications, and protocols for LLM-based autonomous agents into a proposed taxonomy with recommendations for future research.
A cross-disciplinary review of 151 studies concludes LLMs accelerate research workflows while introducing recurring technical and ethical risks, including ten it flags as underexplored.
citing papers explorer
-
SCICONVBENCH: Benchmarking LLMs on Multi-Turn Clarification for Task Formulation in Computational Science
SCICONVBENCH is a new benchmark evaluating LLMs on multi-turn disambiguation and inconsistency resolution for task formulation in computational science, with frontier models reaching only 52.7% success on fluid mechanics disambiguation cases.
-
TS-Reasoner: Domain-Oriented Time Series Inference Agents for Reasoning and Automated Analysis
TS-Reasoner is a domain-oriented agent using LLMs, computational tools, and error feedback for multi-step time series inference, showing better performance than general LLMs on understanding and reasoning benchmarks.
-
ALL-FEM: Agentic Large Language models Fine-tuned for Finite Element Methods
ALL-FEM fine-tunes LLMs on a corpus of verified FEniCS scripts and uses multi-agent workflows to automate finite element code generation, achieving 71.79% success on 39 benchmarks across elasticity, flow, and coupled problems.
-
Evo-Memory: Benchmarking LLM Agent Test-time Learning with Self-Evolving Memory
Evo-Memory is a new streaming benchmark and evaluation framework for self-evolving memory in LLM agents, unifying over ten memory modules and introducing the ReMem pipeline for continual improvement on multi-turn and reasoning datasets.
-
IR-Agent: Expert-Inspired LLM Agents for Structure Elucidation from Infrared Spectra
IR-Agent is a multi-agent LLM framework that emulates expert IR spectral analysis procedures to improve molecular structure elucidation accuracy and adaptability.
-
SciHorizon-DataEVA: An Agentic System for AI-Readiness Evaluation of Heterogeneous Scientific Data
SciHorizon-DataEVA is a multi-agent system that applies Sci-TQA2 principles across four dimensions to assess AI-readiness of heterogeneous scientific data via dynamic profiling and self-correcting evaluation workflows.
-
Perovskite-R1: a domain-specialized large language model for intelligent discovery of precursor additives and experimental design
A fine-tuned LLM called Perovskite-R1, built from curated perovskite literature and material libraries, proposes precursor additives and designs with some experimental validation showing improved stability and performance.
-
From LLM Reasoning to Autonomous AI Agents: A Comprehensive Review
A survey consolidating benchmarks, agent frameworks, real-world applications, and protocols for LLM-based autonomous agents into a proposed taxonomy with recommendations for future research.
-
From Text to Discovery: How Large Language Models Are Reshaping Research Across Scientific and Humanistic Disciplines
A cross-disciplinary review of 151 studies concludes LLMs accelerate research workflows while introducing recurring technical and ethical risks, including ten it flags as underexplored.