Pith. sign in

REVIEW 2 cited by

OmniFusion Technical Report

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2404.06212 v1 pith:GHYYKOU3 submitted 2024-04-09 cs.CV cs.AIcs.LG

classification cs.CVcs.AIcs.LG
keywords omnifusionopen-sourceadaptersdifferentencodingimagemodelpropose
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Last year, multimodal architectures served up a revolution in AI-based approaches and solutions, extending the capabilities of large language models (LLM). We propose an \textit{OmniFusion} model based on a pretrained LLM and adapters for visual modality. We evaluated and compared several architecture design principles for better text and visual data coupling: MLP and transformer adapters, various CLIP ViT-based encoders (SigLIP, InternVIT, etc.), and their fusing approach, image encoding method (whole image or tiles encoding) and two 7B LLMs (the proprietary one and open-source Mistral). Experiments on 8 visual-language benchmarks show the top score for the best OmniFusion setup in terms of different VQA tasks in comparison with open-source LLaVA-like solutions: VizWiz, Pope, MM-Vet, ScienceQA, MMBench, TextVQA, VQAv2, MMMU. We also propose a variety of situations, where OmniFusion provides highly-detailed answers in different domains: housekeeping, sightseeing, culture, medicine, handwritten and scanned equations recognition, etc. Mistral-based OmniFusion model is an open-source solution with weights, training and inference scripts available at https://github.com/AIRI-Institute/OmniFusion.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Addressing Hallucinations in Language Models with Knowledge Graph Embeddings as an Additional Modality

    cs.CL 2024-11 conditional novelty 6.0 of 10

    Injecting Wikidata entity embeddings through a linear adapter into frozen LLMs improves hallucination detection on HaluEval, True-False, and FEVER benchmarks.

  2. Visually Interpretable Subtask Reasoning for Visual Question Answering

    cs.CV 2025-05 conditional novelty 5.0 of 10

    Training a visual question answering model on LLM-generated subtask rationales with bounding boxes improves GQA accuracy from 64.0 to 65.1 percent and adds grounded explanations.

Pith tools