Across nine vision-language models, performance collapses when chemical composition is held out, but the reported magnitude and internal consistency of this collapse are not supported by the paper's own tables.
Title resolution pending
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CV 1years
2025 1verdicts
REJECT 1representative citing papers
citing papers explorer
-
Stress-Testing Multimodal Foundation Models for Crystallographic Reasoning
Across nine vision-language models, performance collapses when chemical composition is held out, but the reported magnitude and internal consistency of this collapse are not supported by the paper's own tables.