DiagFlowBench is a new dataset of 1,676 conversations from industrial flowcharts showing that language models often select contextually wrong but real procedure steps instead of abstaining on out-of-scope inputs.
Flowagent: Achieving compliance and flexibility for workflow agents
4 Pith papers cite this work. Polarity classification is still indexing.
citation-role summary
citation-polarity summary
fields
cs.AI 4years
2026 4roles
background 1polarities
background 1representative citing papers
Mind-ParaWorld creates parallel worlds with atomic facts to evaluate search agents on future scenarios, showing they synthesize evidence well but struggle with collection, coverage, sufficiency judgment, and stopping decisions.
citing papers explorer
-
DiagFlowBench: Evaluating How Language Models Handle Off-Procedure Inputs in Grounded Diagnostic Dialogue
DiagFlowBench is a new dataset of 1,676 conversations from industrial flowcharts showing that language models often select contextually wrong but real procedure steps instead of abstaining on out-of-scope inputs.
-
Evaluating the Search Agent in a Parallel World
Mind-ParaWorld creates parallel worlds with atomic facts to evaluate search agents on future scenarios, showing they synthesize evidence well but struggle with collection, coverage, sufficiency judgment, and stopping decisions.
- Specifying AI-SDLC Processes: A Protocol Language for Human-Agent Boundaries
- Tools as Continuous Flow for Evolving Agentic Reasoning