LGMT is a logic-grounded metamorphic testing framework that detects hidden reasoning defects in LLMs by checking consistency on semantically invariant inputs derived from FOL equivalences.
A survey of large language model agents for question answering
9 Pith papers cite this work. Polarity classification is still indexing.
citation-role summary
citation-polarity summary
years
2026 9roles
background 1polarities
background 1representative citing papers
Introduces InterFLOPBench benchmark and evaluates 14 LLMs on multi-label classification of six floating-point error categories in C code, with top models exceeding 0.88 overall F1 but lower scores on subtle errors like underflow.
DeepDiscovery uses a two-stage framework to improve task-relevant file recovery and downstream performance in large codebases, achieving gains on industrial tasks and SWE-bench Verified.
iPDB adds a predict operator and semantic query optimizations to SQL so that LLM and ML calls run efficiently inside the database, delivering 2.5x average and up to 30x speedup over prior systems.
DCRC applies data-centric methods with adversarial examples and program synthesis to produce verifiable reasoning programs for financial question answering.
ATOM uses a nucleus-electron hierarchy and task-driven RL to generate budget-controllable multi-agent collaboration graphs for LLMs, claiming SOTA performance with up to 30% better token efficiency on six benchmarks.
EngGPT2MoE-16B-A3B matches or exceeds other Italian open-source LLMs on most international benchmarks while remaining competitive on ITALIC, though it trails some top international models.
HEART coordinates role-specialized LLM agents to decompose instructions, validate reachability and constraints, and synthesize executable robotic plans, showing higher success than single-LLM baselines on household tasks.
KnowPilot integrates knowledge retrieval and memory systems into generative agents to achieve better results on domain-specific tasks such as text generation.
citing papers explorer
-
LGMT: Logic-Grounded Metamorphic Testing for Evaluating the Reasoning Reliability of LLMs
LGMT is a logic-grounded metamorphic testing framework that detects hidden reasoning defects in LLMs by checking consistency on semantically invariant inputs derived from FOL equivalences.
-
Benchmarking Large Language Models on Floating-Point Error Classification
Introduces InterFLOPBench benchmark and evaluates 14 LLMs on multi-label classification of six floating-point error categories in C code, with top models exceeding 0.88 overall F1 but lower scores on subtle errors like underflow.
-
From Fragments to Paths: Task-Level Context Recovery for Large Industrial Codebases
DeepDiscovery uses a two-stage framework to improve task-relevant file recovery and downstream performance in large codebases, achieving gains on industrial tasks and SWE-bench Verified.
-
iPDB -- Optimizing Semantic SQL Queries
iPDB adds a predict operator and semantic query optimizations to SQL so that LLM and ML calls run efficiently inside the database, delivering 2.5x average and up to 30x speedup over prior systems.
-
Fighting Numerical Hallucinations via Data-centric Compilation for Online Financial QA
DCRC applies data-centric methods with adversarial examples and program synthesis to produce verifiable reasoning programs for financial question answering.
-
ATOM: Instantiating Budget-Controllable Multi-Agent Collaboration via Nucleus-Electron Hierarchy
ATOM uses a nucleus-electron hierarchy and task-driven RL to generate budget-controllable multi-agent collaboration graphs for LLMs, claiming SOTA performance with up to 30% better token efficiency on six benchmarks.
-
Benchmarking EngGPT2-16B-A3B against Comparable Italian and International Open-source LLMs
EngGPT2MoE-16B-A3B matches or exceeds other Italian open-source LLMs on most international benchmarks while remaining competitive on ITALIC, though it trails some top international models.
-
HEART: Coordination of Heterogeneous Expert Agents for Physically Grounded Robotic Task Planning
HEART coordinates role-specialized LLM agents to decompose instructions, validate reachability and constraints, and synthesize executable robotic plans, showing higher success than single-LLM baselines on household tasks.
-
KnowPilot: Your Knowledge-Driven Copilot for Domain Tasks
KnowPilot integrates knowledge retrieval and memory systems into generative agents to achieve better results on domain-specific tasks such as text generation.