FraudBench shows that current multimodal LLMs and specialized AI-image detectors often fail to spot AI-generated fake damage in refund evidence, with true positive rates frequently below 50% on synthetic subsets while producing false positives on real damage.
Title resolution pending
2 Pith papers cite this work. Polarity classification is still indexing.
2
Pith papers citing it
years
2026 2verdicts
UNVERDICTED 2representative citing papers
AutoRestTest ranked first in fault detection, efficiency, and effectiveness in the SBFT 2026 REST League on 11 APIs with 317 operations under a one-hour budget.
citing papers explorer
-
FraudBench: A Multimodal Benchmark for Detecting AI-Generated Fraudulent Refund Evidence
FraudBench shows that current multimodal LLMs and specialized AI-image detectors often fail to spot AI-generated fake damage in refund evidence, with true positive rates frequently below 50% on synthetic subsets while producing false positives on real damage.
-
AutoRestTest at the SBFT 2026 Tool Competition
AutoRestTest ranked first in fault detection, efficiency, and effectiveness in the SBFT 2026 REST League on 11 APIs with 317 operations under a one-hour budget.