Pith. sign in

Title resolution pending

1 Pith paper cite this work. Polarity classification is still indexing.

1 Pith paper citing it

fields

cs.CV 1

years

2025 1

verdicts

CONDITIONAL 1

representative citing papers

FLIP Reasoning Challenge

cs.CV · 2025-04-16 · conditional · novelty 6.0

The FLIP benchmark of 11,674 blockchain image-story puzzles shows best open and closed AI models reach 75.5% and 77.9% accuracy, below the 95.3% human consensus baseline.

citing papers explorer

Showing 1 of 1 citing paper.

  • FLIP Reasoning Challenge cs.CV · 2025-04-16 · conditional · none · ref 7

    The FLIP benchmark of 11,674 blockchain image-story puzzles shows best open and closed AI models reach 75.5% and 77.9% accuracy, below the 95.3% human consensus baseline.