Pith. sign in

REVIEW 12 cited by

LLaVA-Phi: Efficient Multi-Modal Assistant with Small Language Model

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2401.02330 v4 pith:Y34AZWCY submitted 2024-01-04 cs.CV cs.CL

classification cs.CVcs.CL
keywords multi-modallanguagellava-phimodelmodelsassistantavailabledialogues
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
abstract

In this paper, we introduce LLaVA-$\phi$ (LLaVA-Phi), an efficient multi-modal assistant that harnesses the power of the recently advanced small language model, Phi-2, to facilitate multi-modal dialogues. LLaVA-Phi marks a notable advancement in the realm of compact multi-modal models. It demonstrates that even smaller language models, with as few as 2.7B parameters, can effectively engage in intricate dialogues that integrate both textual and visual elements, provided they are trained with high-quality corpora. Our model delivers commendable performance on publicly available benchmarks that encompass visual comprehension, reasoning, and knowledge-based perception. Beyond its remarkable performance in multi-modal dialogue tasks, our model opens new avenues for applications in time-sensitive environments and systems that require real-time interaction, such as embodied agents. It highlights the potential of smaller language models to achieve sophisticated levels of understanding and interaction, while maintaining greater resource efficiency.The project is available at {https://github.com/zhuyiche/llava-phi}.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 12 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. SCOPE and SCION: A Benchmark and an Auditable Reference Pipeline for Schema Induction and Fusion from Text

    cs.AI 2026-05 conditional novelty 6.0 of 10

    A 24-dataset benchmark for inducing schema graphs from raw text, plus an auditable LLM-based pipeline that reports the highest scores on the benchmark's four schema-similarity metrics.

  2. Boosting Embodied AI Agents through Perception-Generation Disaggregation and Asynchronous Pipeline Execution

    cs.AI 2025-09 conditional novelty 6.0 of 10

    Auras, a perception-generation disaggregation framework with a public context buffer and asynchronous pipeline executor, raises embodied-agent throughput by 2.54x on average without losing accuracy (102.7%).

  3. FUDOKI: Discrete Flow-based Unified Understanding and Generation via Kinetic-Optimal Velocities

    cs.CV 2025-05 conditional novelty 6.0 of 10

    A 1.5B unified multimodal model trained with discrete flow matching and metric-induced probability paths matches autoregressive baselines of similar size on generation and understanding benchmarks.

  4. UniCMs: A Unified Consistency Model For Efficient Multimodal Generation and Understanding

    cs.CV 2025-02 conditional novelty 6.0 of 10

    UniCMs applies consistency distillation to a unified multimodal transformer, treating image mask-diffusion steps and text parallel-decoding steps as one shared denoising trajectory, enabling 2 to 8 step generation and...

  5. WalkVLM:Aid Visually Impaired People Walking by Vision Language Model

    cs.CV 2024-12 conditional novelty 6.0 of 10

    WalkVLM combines chain-of-thought reasoning with a temporal trigger module, and the new Walking Awareness Dataset provides 12,000 annotated walking videos to benchmark AI walking assistance.

  6. Olympus: A Universal Task Router for Computer Vision Tasks

    cs.CV 2024-12 conditional novelty 6.0 of 10

    Olympus is a trained MLLM router that delegates 20 vision tasks to specialist models and supports chain-of-action execution of up to five tasks per instruction.

  7. Understanding Museum Exhibits using Vision-Language Reasoning

    cs.CV 2024-12 conditional novelty 6.0 of 10

    A new 65M-image, 200M-QA dataset for museum exhibits lets fine-tuned vision-language models beat general-purpose VLMs on museum attribute questions, especially on questions requiring background knowledge.

  8. Eve: Efficient Multimodal Vision Language Models with Elastic Visual Experts

    cs.CV 2025-01 conditional novelty 5.0 of 10

    A 1.8B-parameter vision-language model, Eve, uses elastic visual experts and type-aware token routing to reach a 68.87% average on six VLM benchmarks while preserving language performance.

  9. Learn from Downstream and Be Yourself in Multimodal Large Language Model Fine-Tuning

    cs.CL 2024-11 conditional novelty 5.0 of 10

    SPIDER updates only parameters whose fine-tuning gradient importance exceeds their pre-trained weight importance, reducing catastrophic forgetting and improving downstream performance in multimodal LLM fine-tuning.

  10. LLMQuoter: Enhancing RAG Capabilities Through Efficient Quote Extraction From Large Contexts

    cs.CL 2025-01 reject novelty 4.0 of 10

    LLMQuoter uses a distilled 3B model to extract quotes for RAG; the paper shows gold quotes greatly improve QA, but does not test its own model's quotes end-to-end.

  11. AlzheimerRAG: Multimodal Retrieval Augmented Generation for Clinical Use Cases using PubMed articles

    cs.IR 2024-12 reject novelty 3.0 of 10

    AlzheimerRAG, a PubMed-based multimodal retrieval-augmented generation system, is reported, but its PubMedQA results are in-sample because PubMedQA was used for both fine-tuning and testing.

  12. Generalist Virtual Agents: A Survey on Autonomous Agents Across Digital Platforms

    cs.MA 2024-11 conditional novelty 3.0 of 10

    A survey that proposes the Generalist Virtual Agent concept and taxonomies for agent environments, tasks, perceptions, actions, models, and evaluation, concluding that real-world-like environments favor human-like int...

Pith tools