VLP adds an NL documentation layer with trace-linked mismatch detection and derived formal checks to make human validation of LLM code feasible, lifting pass@1 from 28.7-73.2% to 65.4-93.5%.
Llm-check: Investigating detection of hallucinations in large language models
2 Pith papers cite this work. Polarity classification is still indexing.
2
Pith papers citing it
citation-role summary
background 1
citation-polarity summary
years
2026 2verdicts
UNVERDICTED 2roles
background 1polarities
background 1representative citing papers
CAESAR decomposes LLM-based intrusion workflows into five roles with bounded coordination protocols, yielding higher success rates and lower variance than single-agent baselines on 25 CTF tasks.
citing papers explorer
-
Guiding Human Validation of LLM-Generated Code via Verifiable Literate Programming
VLP adds an NL documentation layer with trace-linked mismatch detection and derived formal checks to make human validation of LLM code feasible, lifting pass@1 from 28.7-73.2% to 65.4-93.5%.
-
When LLMs Team Up: A Coordinated Attack Framework for Automated Cyber Intrusions
CAESAR decomposes LLM-based intrusion workflows into five roles with bounded coordination protocols, yielding higher success rates and lower variance than single-agent baselines on 25 CTF tasks.