SF20K competition results show top AI models reach 65.7% accuracy on story-level video QA using amateur films, with information selection and reasoning structure as the main bottlenecks rather than model size, well below human performance of 91.7%.
Introducing GPT-4.1 in the API, 2025
2 Pith papers cite this work. Polarity classification is still indexing.
verdicts
UNVERDICTED 2representative citing papers
MaxShapley computes fair document attributions in generative QA by reducing Shapley value calculation to polynomial time via a max-sum utility, matching exact Shapley quality on HotPotQA, MuSiQUE, and MS MARCO while using up to 9x fewer resources.
citing papers explorer
-
SF20K Competition 2025: Summary and findings
SF20K competition results show top AI models reach 65.7% accuracy on story-level video QA using amateur films, with information selection and reasoning structure as the main bottlenecks rather than model size, well below human performance of 91.7%.
-
MaxShapley: Towards Incentive-compatible Generative Search with Fair Context Attribution
MaxShapley computes fair document attributions in generative QA by reducing Shapley value calculation to polynomial time via a max-sum utility, matching exact Shapley quality on HotPotQA, MuSiQUE, and MS MARCO while using up to 9x fewer resources.