Pith. sign in

REVIEW 13 cited by

ImageBind: One Embedding Space To Bind Them All

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2305.05665 v2 pith:IKQ7JTYS submitted 2023-05-09 cs.CV cs.AIcs.LGcs.MM

ImageBind: One Embedding Space To Bind Them All

classification cs.CV cs.AIcs.LGcs.MM
keywords modalitiesimagebinddataembeddingemergentmodelsacrossbind
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
read the original abstract

We present ImageBind, an approach to learn a joint embedding across six different modalities - images, text, audio, depth, thermal, and IMU data. We show that all combinations of paired data are not necessary to train such a joint embedding, and only image-paired data is sufficient to bind the modalities together. ImageBind can leverage recent large scale vision-language models, and extends their zero-shot capabilities to new modalities just by using their natural pairing with images. It enables novel emergent applications 'out-of-the-box' including cross-modal retrieval, composing modalities with arithmetic, cross-modal detection and generation. The emergent capabilities improve with the strength of the image encoder and we set a new state-of-the-art on emergent zero-shot recognition tasks across modalities, outperforming specialist supervised models. Finally, we show strong few-shot recognition results outperforming prior work, and that ImageBind serves as a new way to evaluate vision models for visual and non-visual tasks.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 13 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Fusion Embedding: A Unified Embedding Space for Text, Image, Video, and Audio

    cs.CL 2026-07 conditional novelty 7.0

    Trained connectors and audio-only gated adapters integrate audio into a frozen vision-language embedding space, preserving base outputs bit-exactly and yielding emergent audio-image retrieval.

  2. When to Align, When to Predict: A Phase Diagram for Multimodal Learning

    cs.LG 2026-06 accept novelty 7.0

    A spiked signal-plus-noise model yields separation ratios that partition multimodal problems into four regimes where alignment, prediction, both, or neither succeed.

  3. Don't Let the Video Speak: Audio-Contrastive Preference Optimization for Audio-Visual Language Models

    cs.CV 2026-04 unverdicted novelty 7.0

    Audio-Contrastive Preference Optimization (ACPO) mitigates audio hallucination in AVLMs via output-contrastive and input-contrastive objectives that enforce faithful audio grounding.

  4. EmergentBridge: Improving Zero-Shot Cross-Modal Transfer in Unified Multimodal Embedding Models

    cs.AI 2026-04 unverdicted novelty 7.0

    EmergentBridge improves zero-shot cross-modal transfer for unpaired modality pairs by learning noisy bridge anchors and enforcing proxy alignment only in the orthogonal subspace to preserve existing anchor alignments.

  5. OmniVAE: An Audio-Video VAE with Cross-Modal Alignment for Joint Generation

    cs.SD 2026-07 conditional novelty 6.0

    A jointly trained audio-video VAE with segment-level contrastive alignment and semantic distillation improves downstream text-to-audio-video synchronization and quality.

  6. OmniVAE: An Audio-Video VAE with Cross-Modal Alignment for Joint Generation

    cs.SD 2026-07 conditional novelty 6.0

    Jointly training audio and video VAEs with segment contrastive loss and semantic distillation yields more learnable, cross-aligned latents that improve downstream joint generation quality and sync.

  7. EmergentBridge: Improving Zero-Shot Cross-Modal Transfer in Unified Multimodal Embedding Models

    cs.AI 2026-04 unverdicted novelty 6.0

    EmergentBridge enhances zero-shot cross-modal performance on unpaired modalities by learning noisy bridge anchors from existing alignments and enforcing proxy alignment only in the orthogonal subspace to avoid gradien...

  8. Echoes Over Time: Unlocking Length Generalization in Video-to-Audio Generation Models

    cs.CV 2026-02 unverdicted novelty 6.0

    MMHNet enables video-to-audio models trained on short clips to generalize and generate audio for videos over 5 minutes long.

  9. Assessing Factual Music Comprehension in Large Audio Language Models

    cs.SD 2025-11 conditional novelty 6.0

    Standard NLP metrics fail to capture factual music understanding in audio-language models; a CLAP-based metric and an LLM-parsed factual QA protocol measure it more directly.

  10. Artificial Phantasia: Emergent Mental Imagery in Large Language Models

    cs.AI 2025-09 unverdicted novelty 6.0

    LLMs achieve higher accuracy than humans on compositional imagery tasks previously argued to require pictorial representations, supporting emergent propositional mental imagery in AI.

  11. Testing chatbots on the creation of encoders for audio conditioned image generation

    cs.SD 2025-09 conditional novelty 6.0

    All chatbot-designed audio encoders failed to align with CLIP text embeddings and produced incoherent images, while showing a surprising architectural similarity across chatbots.

  12. PandaGPT: One Model To Instruction-Follow Them All

    cs.CL 2023-05 conditional novelty 6.0

    A single model trained only on image-text pairs gains instruction-following ability across images, video, and audio by routing all modalities through ImageBind's shared embedding space into Vicuna.

  13. AlphaWiSE: Adaptive Weight Interpolation for Continual Multimodal Representation Learning

    cs.CV 2026-07 conditional novelty 5.0

    Fitting one interpolation coefficient per parameter tensor on a small exemplar memory improves continual audio–image–text retrieval over individual continual-learning checkpoints.