Gemma3 vision-language models can label UI snapshot test failure causes with 84% recall on a small synthetic iOS dataset, but prompt-based selective ignore is unreliable.
An empirical study on the use of snapshot testing,
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.SE 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
LLMShot: Reducing snapshot testing maintenance via LLMs
Gemma3 vision-language models can label UI snapshot test failure causes with 84% recall on a small synthetic iOS dataset, but prompt-based selective ignore is unreliable.