← back to paper
arxiv: 2607.05264 · 2 revisions
SteelBench: Evaluating Vision-Language Models in Real-World Industrial Environments