DEBENCH shows the best decompiler-LLM pair reaches only 22.3% program-level behavioral overlap and 1.2% exact stdout match, with decompiler engines driving 20x more variation than LLMs.
Is your benchmark (still) useful? dynamic benchmarking for code language models
2 Pith papers cite this work. Polarity classification is still indexing.
2
Pith papers citing it
citation-role summary
background 1
citation-polarity summary
years
2026 2verdicts
UNVERDICTED 2roles
background 1polarities
background 1representative citing papers
Agent benchmarks can report evidence-supported score bounds instead of single misleading success rates by adding a layer that checks required artifacts for outcome verification.
citing papers explorer
-
CODEFUSE-DEBENCH: An Empirical Study on Readability, Recompilability, and Functionality
DEBENCH shows the best decompiler-LLM pair reaches only 22.3% program-level behavioral overlap and 1.2% exact stdout match, with decompiler engines driving 20x more variation than LLMs.
-
Can Agent Benchmarks Support Their Scores? Evidence-Supported Bounds for Interactive-Agent Evaluation
Agent benchmarks can report evidence-supported score bounds instead of single misleading success rates by adding a layer that checks required artifacts for outcome verification.