DiagnosticIQ benchmark shows frontier LLMs perform similarly on standard rule-to-action tasks but lose substantial accuracy under distractor expansion and condition inversion, pointing to calibration as the key deployment issue.
Assetopsbench: Benchmarking ai agents for task automation in industrial asset operations and maintenance
12 Pith papers cite this work. Polarity classification is still indexing.
citation-role summary
citation-polarity summary
years
2026 12representative citing papers
PHMForge benchmark shows LLM agents achieve 80.8% pass@1 on prognostic tasks with native MCP tools but performance collapses from 100% to 20% when using text RAG instead.
FacProcessTwin uses an LLM to generate process twin models from documentation and natural language, automatically binds them to operational data, and uses human oversight for safety-critical steps, achieving 95.2% mean F1 accuracy and one-sixth the manual development time in a 16-flow food manufactu
IndustryBench is a standards-grounded Chinese benchmark that exposes LLMs' persistent gaps in industrial terminology, safety compliance, and parameter accuracy, with safety checks reshuffling model rankings.
Aggregate leaderboards for LLM agents lack predictive validity for out-of-distribution settings, and the paper proposes ranking by in-sample to out-of-sample rank correlation instead of mean score.
Trajel introduces a five-type taxonomy and benchmark for trajectory-level hallucinations in multi-agent LLM workflows, showing existing final-answer benchmarks miss common failures.
SPIN enforces DAG-valid plans and prefix-based stopping for LLM agents, cutting executed tasks from 1061 to 623 and tool calls from 11.81 to 6.82 per run on AssetOpsBench while raising success from 0.638 to 0.706.
Typed knowledge graphs improve LLM accuracy on industrial asset operations from 65% to 82-99% via structured retrieval, deterministic execution, and generation-augmented knowledge for missing facts.
QLoRA fine-tuning on tool-use data enables 4B-parameter models to perform structured planning without tool catalogs in prompts, outperforming informed baselines on AssetOpsBench while reducing input length by 82.6%.
Retrospective of a 2025 AI agent competition finds public-private score misalignment, an inert composite component, multi-account registrations, and guardrail fixes outperforming architectural novelty.
A literature survey finds foundation-model agents in industry are 75% at prototype stages with gains in human interaction and uncertainty handling but deficits in negotiation, plus limitations like hallucinations and latency.
DynAMO applies topological scheduling to LLM agent workflows for industrial asset management, claiming 1.6-1.8x latency reduction on AssetOpsBench while preserving safety and showing model inference as the main bottleneck.
citing papers explorer
-
DiagnosticIQ: A Benchmark for LLM-Based Industrial Maintenance Action Recommendation from Symbolic Rules
DiagnosticIQ benchmark shows frontier LLMs perform similarly on standard rule-to-action tasks but lose substantial accuracy under distractor expansion and condition inversion, pointing to calibration as the key deployment issue.
-
PHMForge: Evaluating LLM Agents on Industrial Prognostics through MCP-Native, Algorithm-Grounded Tools
PHMForge benchmark shows LLM agents achieve 80.8% pass@1 on prognostic tasks with native MCP tools but performance collapses from 100% to 20% when using text RAG instead.
-
FacProcessTwin: An LLM-Based System for Process Twin Development
FacProcessTwin uses an LLM to generate process twin models from documentation and natural language, automatically binds them to operational data, and uses human oversight for safety-critical steps, achieving 95.2% mean F1 accuracy and one-sixth the manual development time in a 16-flow food manufactu
-
IndustryBench: Probing the Industrial Knowledge Boundaries of LLMs
IndustryBench is a standards-grounded Chinese benchmark that exposes LLMs' persistent gaps in industrial terminology, safety compliance, and parameter accuracy, with safety checks reshuffling model rankings.
-
Beyond Static Leaderboards: Predictive Validity for the Evaluation of LLM Agents
Aggregate leaderboards for LLM agents lack predictive validity for out-of-distribution settings, and the paper proposes ranking by in-sample to out-of-sample rank correlation instead of mean score.
-
Beyond Final Answers: Auditing Trajectory-Level Hallucinations in Multi-Agent Industrial Workflows
Trajel introduces a five-type taxonomy and benchmark for trajectory-level hallucinations in multi-agent LLM workflows, showing existing final-answer benchmarks miss common failures.
-
SPIN: Structural LLM Planning via Iterative Navigation for Industrial Tasks
SPIN enforces DAG-valid plans and prefix-based stopping for LLM agents, cutting executed tasks from 1061 to 623 and tool calls from 11.81 to 6.82 per run on AssetOpsBench while raising success from 0.638 to 0.706.
-
Knowledge Graphs as the Missing Data Layer for LLM-Based Industrial Asset Operations
Typed knowledge graphs improve LLM accuracy on industrial asset operations from 65% to 82-99% via structured retrieval, deterministic execution, and generation-augmented knowledge for missing facts.
-
Internalizing Tool Knowledge in Small Language Models via QLoRA Fine-Tuning
QLoRA fine-tuning on tool-use data enables 4B-parameter models to perform structured planning without tool catalogs in prompts, outperforming informed baselines on AssetOpsBench while reducing input length by 82.6%.
-
Results and Retrospective Analysis of the CODS 2025 AssetOpsBench Challenge
Retrospective of a 2025 AI agent competition finds public-private score misalignment, an inert composite component, multi-account registrations, and guardrail fixes outperforming architectural novelty.
-
Foundation-Model-Based Agents in Industrial Automation: Purposes, Capabilities, and Open Challenges
A literature survey finds foundation-model agents in industry are 75% at prototype stages with gains in human interaction and uncertainty handling but deficits in negotiation, plus limitations like hallucinations and latency.
-
DynAMO:Dynamic Asset Management Orchestration via Topological Multi-Agent Scheduling
DynAMO applies topological scheduling to LLM agent workflows for industrial asset management, claiming 1.6-1.8x latency reduction on AssetOpsBench while preserving safety and showing model inference as the main bottleneck.