ArchSIBench is a new benchmark dataset and evaluation suite that measures vision-language models on architectural spatial intelligence across 17 subtasks, showing most models lag human baselines especially in transformation and configuration.
Aecv-bench: Benchmarking multimodal models on architectural and engineering drawings understanding.arXiv preprint arXiv:2601.04819
2 Pith papers cite this work. Polarity classification is still indexing.
2
Pith papers citing it
fields
cs.CV 2years
2026 2representative citing papers
MechVQA is a 10-task mechanical-drawing VQA benchmark; MechVL-4B trained with SFT plus DAPO RL scores 84.85, +7.57 over the best closed-source baseline.
citing papers explorer
-
ArchSIBench: Benchmarking the Architectural Spatial Intelligence of Vision-Language Models
ArchSIBench is a new benchmark dataset and evaluation suite that measures vision-language models on architectural spatial intelligence across 17 subtasks, showing most models lag human baselines especially in transformation and configuration.
-
MechVQA: Benchmarking and Enhancing Multimodal LLMs on Comprehensive Mechanical Drawing Understanding
MechVQA is a 10-task mechanical-drawing VQA benchmark; MechVL-4B trained with SFT plus DAPO RL scores 84.85, +7.57 over the best closed-source baseline.