FailSafeQA, a 220-example financial long-context benchmark, shows no tested LLM can both stay robust to input perturbations and refuse to hallucinate when context is missing or irrelevant.
Title resolution pending
1 Pith paper cite this work, alongside 13 external citations. Polarity classification is still indexing.
1
Pith paper citing it
13
external citations · OpenAlex
fields
cs.CL 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Expect the Unexpected: FailSafe Long Context QA for Finance
FailSafeQA, a 220-example financial long-context benchmark, shows no tested LLM can both stay robust to input perturbations and refuse to hallucinate when context is missing or irrelevant.