Pith. sign in

REVIEW 4 cited by

LEGO Co-builder: Exploring Fine-Grained Vision-Language Modeling for Multimodal LEGO Assembly Assistants

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2507.05515 v2 pith:XBT7QP7R submitted 2025-07-07 cs.AI cs.CLcs.CV

LEGO Co-builder: Exploring Fine-Grained Vision-Language Modeling for Multimodal LEGO Assembly Assistants

classification cs.AI cs.CLcs.CV
keywords assemblydetectionfine-grainedlegomultimodalstateassistantsbenchmark
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
read the original abstract

Vision-language models (VLMs) are facing the challenges of understanding and following multimodal assembly instructions, particularly when fine-grained spatial reasoning and precise object state detection are required. In this work, we explore LEGO Co-builder, a hybrid benchmark combining real-world LEGO assembly logic with programmatically generated multimodal scenes. The dataset captures stepwise visual states and procedural instructions, allowing controlled evaluation of instruction-following, object detection, and state detection. We introduce a unified framework and assess leading VLMs such as GPT-4o, Gemini, and Qwen-VL, under zero-shot and fine-tuned settings. Our results reveal that even advanced models like GPT-4o struggle with fine-grained assembly tasks, with a maximum F1 score of just 40.54\% on state detection, highlighting gaps in fine-grained visual understanding. We release the benchmark, codebase, and generation pipeline to support future research on multimodal assembly assistants grounded in real-world workflows.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. AssemblyBench: Physics-Aware Assembly of Complex Industrial Objects

    cs.CV 2026-05 unverdicted novelty 7.0

    AssemblyBench dataset and AssemblyDyno transformer model enable physics-aware prediction of assembly sequences and trajectories for complex industrial objects from multimodal instructions and 3D shapes.

  2. Brick-Composer: Using MLLMs for Assembly with Diverse Bricks

    cs.AI 2026-06 unverdicted novelty 6.0

    Brick-Composer trains MLLMs on brick assembly via three signals, raising step-level success from under 1% to around 15% on the new BC-Bench benchmark.

  3. Operating Within the Operational Design Domain: Zero-Shot Perception with Vision-Language Models

    cs.CV 2026-05 unverdicted novelty 5.0

    Vision-language models achieve usable zero-shot ODD perception in driving scenes when guided by definition-anchored chain-of-thought prompting with persona decomposition.

  4. Operating Within the Operational Design Domain: Zero-Shot Perception with Vision-Language Models

    cs.CV 2026-05 unverdicted novelty 4.0

    Vision-language models can serve as zero-shot ODD sensors for autonomous driving when using definition-anchored chain-of-thought prompting with persona decomposition.