Static CTF benchmarks for LLM agents are contamination-prone; CTFusion evaluates agents on live CTFs via an MCP server on CTFd with per-agent isolation.
LiveBench: A Challenging, Contamination-Free
2 Pith papers cite this work. Polarity classification is still indexing.
2
Pith papers citing it
fields
cs.LG 2years
2026 2verdicts
UNVERDICTED 2representative citing papers
Generative AI evaluation must shift from static benchmark scores to measuring sustained improvements in human capabilities within specific deployment contexts.
citing papers explorer
-
CTFusion: A CTF-based Benchmark for LLM Agent Evaluation
Static CTF benchmarks for LLM agents are contamination-prone; CTFusion evaluates agents on live CTFs via an MCP server on CTFd with per-agent isolation.
-
Benchmarked Yet Not Measured -- Generative AI Should be Evaluated Against Real-World Utility
Generative AI evaluation must shift from static benchmark scores to measuring sustained improvements in human capabilities within specific deployment contexts.