Pith. sign in

REVIEW 9 cited by

LLaVA-Interactive: An All-in-One Demo for Image Chat, Segmentation, Generation and Editing

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2311.00571 v1 pith:O3B7VZWY submitted 2023-11-01 cs.CV cs.AIcs.CLcs.HCcs.MM

LLaVA-Interactive: An All-in-One Demo for Image Chat, Segmentation, Generation and Editing

classification cs.CV cs.AIcs.CLcs.HCcs.MM
keywords llava-interactivemultimodalimagechateditinggenerationhumaninteraction
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
Share X Bluesky LinkedIn Reddit HN
read the original abstract

LLaVA-Interactive is a research prototype for multimodal human-AI interaction. The system can have multi-turn dialogues with human users by taking multimodal user inputs and generating multimodal responses. Importantly, LLaVA-Interactive goes beyond language prompt, where visual prompt is enabled to align human intents in the interaction. The development of LLaVA-Interactive is extremely cost-efficient as the system combines three multimodal skills of pre-built AI models without additional model training: visual chat of LLaVA, image segmentation from SEEM, as well as image generation and editing from GLIGEN. A diverse set of application scenarios is presented to demonstrate the promises of LLaVA-Interactive and to inspire future research in multimodal interactive systems.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 9 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. A Comprehensive Study of Implementation Bugs in Multi-modal Agents

    cs.SE 2026-07 accept novelty 7.0

    First systematic taxonomy of 158 multi-modal agent bugs plus a runtime analyzer that recovers most open issues and surfaces 31 new ones.

  2. Don't Guess, Just Ask: Resolving Ambiguity in Referring Segmentation via Multi-turn Clarification

    cs.CV 2026-05 unverdicted novelty 7.0

    IC-Seg is a new agentic framework using multi-turn clarification and Hi-GRPO hierarchical optimization to resolve ambiguous queries in referring video object segmentation while maintaining performance on standard benchmarks.

  3. Relightable Gaussian Splatting for Virtual Production Using Image-Based Illumination

    cs.CV 2026-05 unverdicted novelty 7.0

    A relightable Gaussian Splatting method for virtual production decomposes scenes into fixed appearance and variable lighting by parameterizing primitives to directly sample high-resolution background textures, enablin...

  4. Don't Guess, Just Ask: Resolving Ambiguity in Referring Segmentation via Multi-turn Clarification

    cs.CV 2026-05 unverdicted novelty 6.0

    IC-Seg is a multi-turn clarification framework with hierarchical GRPO optimization that resolves ambiguous queries in referring video object segmentation and introduces the Ambi-RVOS benchmark.

  5. Ego-centric Predictive Model Conditioned on Hand Trajectories

    cs.CV 2025-08 conditional novelty 6.0

    Ego-PM predicts future hand trajectories and then uses them to condition latent diffusion video generation, jointly outputting actions and future frames in egocentric and robotic scenes.

  6. Training-Free Multimodal Large Language Model Orchestration

    cs.CL 2025-08 unverdicted novelty 6.0

    LLM Orchestration integrates modality experts via an LLM controller, cross-modal memory, and interaction layer to enable multimodal input-output without gradient-based training.

  7. Sa2VA: Marrying SAM2 with LLaVA for Dense Grounded Understanding of Images and Videos

    cs.CV 2025-01 conditional novelty 6.0

    Sa2VA unifies SAM-2 segmentation with MLLM reasoning into a single model for referring segmentation and conversation on images and videos, supported by a new 72k-expression Ref-SAV dataset.

  8. Training-Free Multimodal Large Language Model Orchestration

    cs.CL 2025-08 unverdicted novelty 5.0

    A training-free orchestration framework integrates off-the-shelf modality experts via an LLM controller, text-centric cross-modal memory, and unified interaction layer to enable multimodal input-output without joint training.

  9. Document Parsing Unveiled: Techniques, Challenges, and Prospects for Structured Information Extraction

    cs.MM 2024-10 unverdicted novelty 3.0

    Survey proposing a taxonomy for document parsing into pipeline-based systems and VLM-driven unified models, reviewing components, metrics, benchmarks, and challenges.