Pith. sign in

REVIEW 4 cited by

Visual Programming: Compositional visual reasoning without training

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2211.11559 v1 pith:7UF5AC5A submitted 2022-11-18 cs.CV cs.AIcs.CL

classification cs.CVcs.AIcs.CL
keywords visprogvisualcompositionalimagetaskscomplexlanguagemodels
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We present VISPROG, a neuro-symbolic approach to solving complex and compositional visual tasks given natural language instructions. VISPROG avoids the need for any task-specific training. Instead, it uses the in-context learning ability of large language models to generate python-like modular programs, which are then executed to get both the solution and a comprehensive and interpretable rationale. Each line of the generated program may invoke one of several off-the-shelf computer vision models, image processing routines, or python functions to produce intermediate outputs that may be consumed by subsequent parts of the program. We demonstrate the flexibility of VISPROG on 4 diverse tasks - compositional visual question answering, zero-shot reasoning on image pairs, factual knowledge object tagging, and language-guided image editing. We believe neuro-symbolic approaches like VISPROG are an exciting avenue to easily and effectively expand the scope of AI systems to serve the long tail of complex tasks that people may wish to perform.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. CanvasAgent: Enabling Complex Image Creation and Editing via Visual Tool Orchestration

    cs.CV 2026-07 conditional novelty 6.0 of 10

    SFT+GRPO training on CanvasCraft teaches an MLLM to orchestrate heterogeneous visual tools for long-horizon image creation and editing.

  2. Reinforced Visual Perception with Tools

    cs.CV 2025-09 conditional novelty 5.0 of 10

    ReVPT uses GRPO reinforcement learning with a cold-start SFT phase to make Qwen2.5-VL models call visual tools, improving perception benchmarks over SFT and text-only RL baselines.

  3. Multimodal Video Emotion Recognition with Reliable Reasoning Priors

    cs.CV 2025-07 conditional novelty 5.0 of 10

    Injecting MLLM-generated multimodal reasoning traces into a fused audio-visual-text emotion model, plus a balanced dual-contrastive loss, raises average accuracy on MER2024 from 77.5 to 84.7 percent.

  4. Augmented Vision-Language Models: A Systematic Review

    cs.CL 2025-07 conditional novelty 5.0 of 10

    A structured taxonomy of inference-time augmentation techniques that connect vision-language models to external symbolic systems, tools, and knowledge sources.

Pith tools