Pith. sign in

REVIEW 7 cited by

VALOR: Vision-Audio-Language Omni-Perception Pretraining Model and Dataset

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2304.08345 v2 pith:6YZWALWP submitted 2023-04-17 cs.LG cs.CLcs.CVcs.MMeess.AS

classification cs.LGcs.CLcs.CVcs.MMeess.AS
keywords valormultimodalpretrainingaudiomodelvisionvision-audio-languagevision-language
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

In this paper, we propose a Vision-Audio-Language Omni-peRception pretraining model (VALOR) for multi-modal understanding and generation. Different from widely-studied vision-language pretraining models, VALOR jointly models relationships of vision, audio and language in an end-to-end manner. It contains three separate encoders for single modality representations, and a decoder for multimodal conditional text generation. We design two pretext tasks to pretrain VALOR model, including Multimodal Grouping Alignment (MGA) and Multimodal Grouping Captioning (MGC). MGA projects vision, language and audio to the same common space, building vision-language, audio-language and audiovisual-language alignment simultaneously. MGC learns how to generate text tokens in conditions of vision, audio or their both. To promote vision-audio-language pretraining research, we construct a large-scale high-quality tri-modality dataset named VALOR-1M, which contains 1M audiable videos with human annotated audiovisual captions. Extensive experiments show that VALOR can learn strong multimodal correlations and be generalized to various downstream tasks (e.g., retrieval, captioning and question answering), with different input modalities (e.g., vision-language, audio-language and audiovisual-language). VALOR achieves new state-of-the-art performances on series of public cross-modality benchmarks. Code and data are available at project page https://casia-iva-group.github.io/projects/VALOR.

Discussion (0). Sign in to comment.

Forward citations

Cited by 7 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. A Benchmark for Omni-Modal Reasoning in Long Videos

    cs.CV 2025-12 reject novelty 7.0 of 10

    A new 45-minute-scale omni-modal video Q&A benchmark and a training-free retrieval-refine agent, whose reported agent score (66.64% in the abstract) is not supported by the paper's own main results (44.66%).

  2. Empowering Long-form Omni-modal Understanding with Robust Audio Perception

    cs.LG 2026-07 conditional novelty 6.0 of 10

    Decoupled audio-visual caption and CoT-QA datasets plus two-stage fine-tuning measurably strengthen auditory perception and cross-modal reasoning in a 7B omni-modal LLM.

  3. EchoingPixels: Aliasing-Resistant Joint Token Reduction for Audio-Visual LLMs

    cs.CV 2025-12 conditional novelty 6.0 of 10

    EchoingPixels prunes audio-visual LLM input tokens jointly across modalities and re-tunes RoPE frequencies so that 5–20% of tokens retain roughly full-model performance.

  4. CAT-SG: A Large Dynamic Scene Graph Dataset for Fine-Grained Understanding of Cataract Surgery

    cs.CV 2025-06 conditional novelty 6.0 of 10

    CAT-SG is a new cataract surgery scene graph dataset with 1.811 million relation annotations, a two-class technique recognition task, and a query-based scene graph generation baseline.

  5. AV-Master: Dual-Path Comprehensive Perception Makes Better Audio-Visual Question Answering

    cs.CV 2025-10 unverdicted novelty 5.0 of 10

    AV-Master reports state-of-the-art accuracy on four audio-visual question answering benchmarks by combining sequential question-guided focus sampling with modality-preference activation.

  6. Open-set Cross Modal Generalization via Multimodal Unified Representation

    cs.CV 2025-07 conditional novelty 5.0 of 10

    The authors propose OSCMG, an open-set version of Cross Modal Generalization, and show their MICU method with masked contrastive learning and unified jigsaw puzzles outperforms prior methods.

  7. From Vision To Language through Graph of Events in Space and Time: An Explainable Self-supervised Approach

    cs.CV 2025-07 conditional novelty 5.0 of 10

    The paper presents GEST, an event-graph representation of videos that is converted automatically into natural language and is also used as a teacher to pre-train end-to-end video captioning models.

Pith tools