GraphInstruct introduces a six-level progressive benchmark with 800 instructions and 1,582 references to diagnose LLM graph generation gaps, plus a verification-guided iterative prompting framework that improves performance.
International Conference on Learning Representations (ICLR) , year =
3 Pith papers cite this work. Polarity classification is still indexing.
years
2026 3representative citing papers
Increasing LLM coding agents' reasoning effort raises cost and process complexity but does not reliably improve model quality across 140 controlled runs on networked anagram game data.
Larger LLMs hallucinate more often despite having the correct concept available because instruction tuning causes probability mass to disperse across alternative surface forms instead of concentrating on one.
citing papers explorer
-
GraphInstruct: A Progressive Benchmark for Diagnosing Capability Gaps in LLM Graph Generation
GraphInstruct introduces a six-level progressive benchmark with 800 instructions and 1,582 references to diagnose LLM graph generation gaps, plus a verification-guided iterative prompting framework that improves performance.
-
An Experimental Design Approach to Evaluating Agentic AI's Autonomous Model Discovery
Increasing LLM coding agents' reasoning effort raises cost and process complexity but does not reliably improve model quality across 140 controlled runs on networked anagram game data.
-
Hallucination as Commitment Failure: Larger LLMs Misfire Despite Knowing the Answer
Larger LLMs hallucinate more often despite having the correct concept available because instruction tuning causes probability mass to disperse across alternative surface forms instead of concentrating on one.