Pith. sign in

REVIEW 2 cited by

InstructSeq: Unifying Vision Tasks with Instruction-conditioned Multi-modal Sequence Generation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2311.18835 v1 pith:6RZ4ZTRA submitted 2023-11-30 cs.CV

classification cs.CV
keywords instructseqinstructionslanguagenaturaltasksvisualflexiblevision
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Empowering models to dynamically accomplish tasks specified through natural language instructions represents a promising path toward more capable and general artificial intelligence. In this work, we introduce InstructSeq, an instruction-conditioned multi-modal modeling framework that unifies diverse vision tasks through flexible natural language control and handling of both visual and textual data. InstructSeq employs a multimodal transformer architecture encompassing visual, language, and sequential modeling. We utilize a visual encoder to extract image features and a text encoder to encode instructions. An autoregressive transformer fuses the representations and generates sequential task outputs. By training with LLM-generated natural language instructions, InstructSeq acquires a strong comprehension of free-form instructions for specifying visual tasks. This provides an intuitive interface for directing capabilities using flexible natural instructions. Without any task-specific tuning, InstructSeq achieves compelling performance on semantic segmentation, referring expression segmentation/comprehension, and image captioning. The flexible control and multi-task unification empower the model with more human-like versatility and generalizability for computer vision. The code will be released soon at https://github.com/rongyaofang/InstructSeq.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Adaptive Classifier-Free Guidance via Dynamic Low-Confidence Masking

    cs.CL 2025-05 conditional novelty 6.0 of 10

    Adaptive Classifier-Free Guidance (A-CFG) re-masks low-confidence tokens in the unconditional input at each generation step, improving reasoning and planning accuracy for masked diffusion language models.

  2. Progressive Scaling Visual Object Tracking

    cs.CV 2025-05 reject novelty 6.0 of 10

    A progressive scaling training strategy with small-teacher distillation and masked-input alignment improves tracking accuracy and powers a new 12-dataset benchmark.

Pith tools