TTP-R1, a retrieval-then-select system fine-tuned with reinforcement learning, reports state-of-the-art average F1 on four multi-label MITRE ATT&CK extraction benchmarks, with sub-technique F1 7.4 points above Claude Sonnet 4.5 with RAG and 28x lower latency.
Title resolution pending
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CR 1years
2026 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Retrieval-Constrained Policy Optimization for Attack Technique Extraction from Cyber Threat Intelligence
TTP-R1, a retrieval-then-select system fine-tuned with reinforcement learning, reports state-of-the-art average F1 on four multi-label MITRE ATT&CK extraction benchmarks, with sub-technique F1 7.4 points above Claude Sonnet 4.5 with RAG and 28x lower latency.