REVIEW 3 major objections 6 minor 1 cited by
Effectively obtaining acoustic, visual and textual data from videos
T0 review · 3 major / 6 minor · reviewed 2026-08-05 · deepseek-v4-flash
Pith's one-line read A six-step, largely automated pipeline converts raw recorded videos into millions of audio-image-text triples, with the paper's own run producing 2,240,231 pairs from 282,081 videos.
desk verdict A large, transparently-built audio-image-text dataset with a reproducible pipeline, but the 'robust semantic connection' claim is unvalidated; worth serious review, not a pass. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The mechanism that carries the argument is the middle-frame extraction heuristic: for each one-second audio segment cut from the video, take the frame that falls approximately at the midpoint of that interval as the segment's visual counterpart. This single rule converts raw temporal co-occurrence into the claimed strict semantic connection between audio and image. Around it, the pipeline adds protective filters (cropping black borders, splitting at abrupt pixel changes, discarding segments with 0.5 seconds of continuous silence or overly dark frames, keeping one pair in three for diversity), a fixed storage format (16kHz, 16-bit, mono, one-second WAV audio; 512x512 RGB24 JPG/PNG images), an
What would settle it
Take a random sample of a few hundred released pairs; for each, play the audio and show a human annotator the paired frame alongside three frames drawn from other pairs, and count how often the correct frame is chosen. Chance is 25 percent; accuracy near chance would mean the middle-frame heuristic is not delivering the claimed semantic connection. A computational analogue measures a pretrained audio-visual agreement score for true pairs and for the same frames paired with audio shifted in time from the same video.
Extended reading notes
Core claim
On the paper's own terms, the central claim is that a fully specified, easily automated pipeline can turn ordinary videos into audio-image-text observations whose three modalities describe the same situation, and that it scales: 282,081 public videos from MUSIC, AudioSetZSL and SoundNet yielded 2,240,231 audio-image pairs, each later paired with a BLIP-generated caption and published as open datasets. The defining move is to treat temporal co-occurrence as alignment: cut the video into one-second segments, take each segment's middle frame, discard segments failing silence, darkness, or scene-cut filters, subsample for diversity, caption each surviving image. Support comes from an Acoustic Di
Load-bearing premise
The method assumes that whatever is audible in a one-second clip of a video is also what is visible in the frame taken at the middle of that second, so that simply cutting the video at that instant creates a true audio-visual match; no check verifies that the sound and the picture actually correspond.
Editorial extensions
If this is right
- Any research group can replicate the pipeline on its own video collection and obtain a custom audio-image-text dataset without manual annotation or specialized hardware; the paper's captioning run was done on a laptop.
- The released corpus (over 2.2 million pairs with captions, in multiple public datasets) gives the community training and evaluation material for tasks that currently lack data: audio-conditioned image generation, audio-to-text captioning, and any-to-any multimodal models.
- Datasets built this way inherit the character of their source videos: continuous unedited recordings from MUSIC, AudioSetZSL and SoundNet yield mostly non-artificial, non-abstract audio-visual relations, which the authors argue simplifies training and convergence.
- Because text is generated from the image, caption failures are localized to image quality; switching to a stronger image-to-text model or ensembling several captioners can upgrade the textual modality without re-running audio-image extraction.
- The reported yield (about 8 pairs per video after filtering and one-in-three subsampling) is a concrete benchmark figure that future extraction pipelines can compare against.
Reading between the lines
- The alignment assumption is testable and is the first thing to check: a retrieval-style probe on the released pairs (does each audio clip identify its own frame among distractors?) would quantify how much semantic signal the middle-frame heuristic preserves; the paper itself only runs silence, darkness, and scene-cut filters, not this check.
- The paper's minimize-modality-conversions principle has a corollary the authors only gesture at: if audio-to-image models are trained on this corpus, their generations could be evaluated by whether they recreate the source video's frames from the audio alone, turning the dataset into a closed-loop benchmark.
- The one-second window is a design choice, not a necessity: sweeping the window length (e.g., 0.5s versus 2s) would map the trade-off between temporal alignment precision and semantic completeness, and could adapt the recipe to other video types such as tutorials, dialogue, or music.
- The pipeline generalizes beyond video: any temporally co-occurring sensor pair (egocentric video with inertial or physiological signals, for instance) could use the same co-occurrence-as-alignment trick to generate paired data at scale, extending the paper's reach well beyond audio-image-text.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes a three-phase pipeline for constructing large audio-image-text datasets from videos: selecting suitable videos, extracting 1-second audio clips paired with the middle frame from each second, and generating textual captions from the frames with BLIP. The authors apply the pipeline to 282,081 videos from MUSIC, AudioSetZSL, and SoundNet, yielding 2,240,231 audio-image pairs and, after captioning, a set of audio-image-text triples released on Kaggle. The paper's central claim is that the extraction 'ensures a robust semantic connection between modalities,' making the dataset useful for training and evaluating multimodal models, particularly audio-conditioned image-to-image generation.
Significance. If the claimed robust semantic connection holds, the contribution is a large, openly released multimodal resource that addresses a genuine gap: existing datasets rarely combine audio, image, and text modalities at scale. The pipeline is described with concrete, reproducible thresholds (black-border threshold, scene-cut threshold, silence duration, subsampling factor, BLIP configuration), and the data are publicly available. The authors also provide useful summary statistics and an Acoustic Diversity Index estimate. These are real strengths. However, the manuscript's central quality claim is not validated by any direct measurement of audio-image or audio-text correspondence, which limits the confidence one can place in the dataset's utility for the stated downstream tasks.
major comments (3)
- [§4.2, steps 3–4; abstract] The paper's central claim—'robust semantic connection between modalities'—rests entirely on temporal co-occurrence: a 1-second audio segment is paired with the frame approximately at its middle, after splitting on scene cuts and filtering silence/dark frames. No quantitative or human evaluation is provided to show that the middle frame is semantically representative of the audio content, or that the two are aligned beyond being drawn from the same raw footage. Given that the stated purpose of the dataset is to train/evaluate audio-conditioned image models, this is a load-bearing gap. I recommend adding a validation study, e.g., a forced-choice human test on a random sample of pairs (audio plus one correct and one incorrect frame from the same video/different video), or a retrieval-style metric (e.g., CLAP/ImageBind embeddings) demonstrating that paired audio and images are more similar t
- [§4.3 and Figure 8] The text modality is generated by BLIP from the image alone, not from the audio. The paper itself acknowledges in Figure 8 that captions are sometimes wrong, and in Section 5 admits there is 'room for improvement' in the filters. Since the text is derived from the image, any audio-image mismatch automatically corrupts the audio-text link as well. The quantitative text statistics (word counts, frequency table) describe surface properties, not semantic correctness. The authors should either present a targeted evaluation of text-image and text-audio alignment (even a modest human-annotated subset) or clearly reposition the dataset as 'weakly aligned' rather than 'robustly semantically connected.'
- [§5, ADI and waveform analysis] The two quantitative tests reported—the aggregate waveform plot and the Acoustic Diversity Index—measure global statistical properties of the audio collection, not the semantic correspondence between paired modalities. The ADI value of ~3.05 indicates diversity across the whole audio set, and the Gaussian-like amplitude distribution is consistent with generic audio; neither says anything about whether a given 1-second clip matches the paired frame. These analyses therefore do not address the central alignment claim and should not be presented as evidence of dataset quality in that regard.
minor comments (6)
- [Abstract and §1] The phrase 'more than 2,000,000 audio-image pairs' appears in the abstract, while Section 5 reports 2,240,231. The counts are consistent, but the abstract's wording could be made exact. Also, 'audio-image-text observations' is used in the abstract, but the text is generated only from the image; the authors may wish to clarify this dependency early.
- [§4, introductory sentence] Typo: 'In this this section' should be 'In this section.'
- [§5, comparison paragraph] Typo: 'researchers need to arduousness search' should likely be 'researchers need to arduously search.'
- [§4.2, step 2] The threshold of 90 for 'abrupt change' is presented without a sensitivity analysis or reference; since this threshold controls scene-cut detection and thus the duration of each fragment, a brief justification or a check on a few videos would strengthen reproducibility.
- [§4.2, step 4] The silence threshold (absolute sample value < 100 for 0.5 s) and dark-frame threshold (mean pixel intensity < 10) are stated as fixed values. It would be helpful to note whether these were tuned or adopted from prior work, and how sensitive the final dataset size is to these choices.
- [§5, Table 3] For the comparison with AudioSetZSL and SoundNet, the table lists audio and video as present, but the 'Ours' row lists audio, image, and text. Since the underlying video source for our dataset is not distributed, it may be clearer to mark 'V' as not directly included in the released Kaggle datasets.
Circularity Check
No significant circularity: the dataset pipeline is self-contained; the semantic-connection claim rests on an unvalidated heuristic assumption, not on a fitted or self-referential derivation.
full rationale
The paper's derivation chain is a data-construction pipeline: select continuous videos (§4.1), split on scene cuts, pair 1-second audio segments with middle frames (§4.2), filter silence/dark frames, subsample, and generate captions with BLIP (§4.3). No parameter is fitted to a subset of the output and then 'predicted' on a closely related quantity; the reported 2,240,231 audio-image pairs are the direct result of applying the stated rules, not a model prediction. The middle-frame heuristic is justified by external citations ([65,38,93,132]) and is not a self-citation; BLIP is an external image-to-text model; the only self-citations are to the released datasets themselves, which are outputs, not load-bearing inputs. The paper's claim that the approach 'ensures a robust semantic connection between modalities' is supported only by the assumption that temporal co-occurrence in a continuous video implies semantic correspondence. That is an empirical proxy assumption, and the paper does not quantitatively validate it; indeed it acknowledges remaining quality issues in Figure 8 and §5. However, an unsupported heuristic is not circularity under the criteria here: the claim is not made true by definition of the extraction rule, nor is any 'prediction' equivalent to its fitted inputs, nor does the argument reduce to a self-citation chain. The acknowledged limitations are correctness/quality concerns, not circularity. The derivation is therefore self-contained, with no circular step to flag.
Assumptions & free parameters
free parameters (7)
- abrupt change detection threshold =
90
- black border intensity threshold =
15
- silence detection amplitude threshold =
100
- continuous silence duration =
0.5 s
- mean pixel intensity threshold =
10
- skip factor =
1/3
- BLIP beam/token settings =
2 beams, min 10, max 20 tokens
assumptions (5)
- domain assumption The middle frame of a 1-second video segment is semantically representative of the segment's content.
- domain assumption Audio and visual content co-occurring in a raw video are semantically related.
- domain assumption BLIP image captioning produces descriptions that adequately represent the image content.
- domain assumption The source datasets (MUSIC, AudioSetZSL, SoundNet) contain continuous, high-quality recordings without copyright conflicts.
- domain assumption The waveform distribution approximating a Gaussian indicates no bias in the audio data.
Cite this review
Pith. "Pith review of Effectively obtaining acoustic, visual and textual data from videos." pith.science (2026). https://pith.science/paper/WRKWD2RU
@misc{pith2026250905786,
author = {Pith},
title = {Pith review of: Effectively obtaining acoustic, visual and textual data from videos},
year = {2026},
howpublished = {\url{https://pith.science/paper/WRKWD2RU}},
note = {Machine review of arXiv:2509.05786}
}
read the original abstract
The increasing use of machine learning models has amplified the demand for high-quality, large-scale multimodal datasets. However, the availability of such datasets, especially those combining acoustic, visual and textual data, remains limited. This paper addresses this gap by proposing a method to extract related audio-image-text observations from videos. We detail the process of selecting suitable videos, extracting relevant data pairs, and generating descriptive texts using image-to-text models. Our approach ensures a robust semantic connection between modalities, enhancing the utility of the created datasets for various applications. We also discuss the challenges encountered and propose solutions to improve data quality. The resulting datasets, publicly available, aim to support and advance research in multimodal data analysis and machine learning.
Figures
Figures from the paper (7 more)
Forward citations
Cited by 1 Pith paper
-
Testing chatbots on the creation of encoders for audio conditioned image generation
All chatbot-designed audio encoders failed to align with CLIP text embeddings and produced incoherent images, while showing a surprising architectural similarity across chatbots.
Reference graph
Works this paper leans on
-
[1]
Andrea Agostinelli, Timo I. Denk, Zal´ an Borsos, Jesse Engel, Mauro Verzetti, Antoine Caillon, Qingqing Huang, Aren Jansen, Adam Roberts, Marco Tagliasacchi, Matt Sharifi, Neil Zeghidour, and Christian Frank. MusicLM: Generating Music From Text. ArXiv, 2301.11325, 2023
arXiv 2023
-
[2]
Don’t Just Assume; Look and Answer: Overcoming Priors for Visual Question Answering
Aishwarya Agrawal, Dhruv Batra, Devi Parikh, and Aniruddha Kembhavi. Don’t Just Assume; Look and Answer: Overcoming Priors for Visual Question Answering. InPro- ceedings of the 2018 IEEE Conference on Computer Vision and Pattern Recognition, pages 4971–4980, 2018
2018
-
[3]
Mistral Models, 2024
Mistral AI. Mistral Models, 2024
2024
-
[4]
Transcripter- Generation of the transcript from audio to text using Deep Learning.International Journal of Computer Sciences and Engineering, 7(1):770–773, 2019
Fatima Ansari, Ramsakal Gupta, Uday Singh, and Fahimur Shaikh. Transcripter- Generation of the transcript from audio to text using Deep Learning.International Journal of Computer Sciences and Engineering, 7(1):770–773, 2019
2019
-
[5]
The Claude 3 Model Family: Opus, Sonnet, Haiku, 2024
Anthropic. The Claude 3 Model Family: Opus, Sonnet, Haiku, 2024
2024
-
[6]
SoundNet: Learning Sound Repre- sentations from Unlabeled Video
Yusuf Aytar, Carl Vondrick, and Antonio Torralba. SoundNet: Learning Sound Repre- sentations from Unlabeled Video. InProceedings of the 30th International Conference on Neural Information Processing Systems, page 892–900, 2016
2016
-
[7]
Multimodal Language Analysis in the Wild: CMU-MOSEI Dataset and Interpretable Dynamic Fusion Graph
AmirAli Bagher Zadeh, Paul Pu Liang, Soujanya Poria, Erik Cambria, and Louis- Philippe Morency. Multimodal Language Analysis in the Wild: CMU-MOSEI Dataset and Interpretable Dynamic Fusion Graph. InProceedings of the 56th Annual Meeting of the Association for Computational Linguistics, pages 2236–2246, 2018
2018
-
[8]
AudioSetCaps: An Enriched Audio-Caption Dataset using Auto- mated Generation Pipeline with Large Audio and Language Models
Jisheng Bai, Haohe Liu, Mou Wang, Dongyuan Shi, Mark Plumbley, Woon-Seng Gan, and Jianfeng Chen. AudioSetCaps: An Enriched Audio-Caption Dataset using Auto- mated Generation Pipeline with Large Audio and Language Models. InAudio Imag- ination: NeurIPS 2024 Workshop AI-Driven Speech, Music, and Sound Generation, 2024
2024
Show all 139 references
-
[9]
Frozen in Time: A Joint Video and Image Encoder for End-to-End Retrieval
Max Bain, Arsha Nagrani, Gul Varol, and Andrew Zisserman. Frozen in Time: A Joint Video and Image Encoder for End-to-End Retrieval. InProceedings of the 2021 IEEE International Conference on Computer Vision, pages 1708–1718, 2021
2021
-
[10]
Are Mod- els Biased on Text without Gender-related Language? InProceedings of the 12th International Conference on Learning Representations, 2024
Catarina G Bel´ em, Preethi Seshadri, Yasaman Razeghi, and Sameer Singh. Are Mod- els Biased on Text without Gender-related Language? InProceedings of the 12th International Conference on Learning Representations, 2024
2024
-
[11]
Ballester
Marcelo Bertalm ´ ıo, Guillermo Sapiro, Vicent Caselles, and C. Ballester. Image in- painting. InProceedings of the 27th Internationl Conference on Computer Graphics and Interactive Techniques Conference, pages 417–424, 2000. 20
2000
-
[12]
Improving Image Generation with Better Captions
James Betker, Gabriel Goh, Li Jing, Tim Brooks, Jianfeng Wang, Linjie Li, Long Ouyang, Juntang Zhuang, Joyce Lee, Yufei Guo, Wesam Manassra, Prafulla Dhariwal, Casey Chu, Yunxin Jiao, and Aditya Ramesh. Improving Image Generation with Better Captions. 2023
2023
-
[13]
Clifton, Yuxiong He, Dacheng Tao, and Shuaiwen Leon Song
Fengxiang Bie, Yibo Yang, Zhongzhu Zhou, Adam Ghanem, Minjia Zhang, Zhewei Yao, Xiaoxia Wu, Connor Holmes, Pareesa Golnari, David A. Clifton, Yuxiong He, Dacheng Tao, and Shuaiwen Leon Song. RenAIssance: A Survey into AI Text-to- Image Generation in the Era of Large Model.ArXi...
2023 arXiv
-
[14]
Birhane and V
A. Birhane and V. Prabhu. Large image datasets: A pyrrhic win for computer vision? InProceedings of the 2021 IEEE Winter Conference on Applications of Computer Vision, pages 1536–1546, 2021
2021
-
[15]
R. B. Blackman and J. W. Tukey. The measurement of power spectra from the point of view of communications engineering — Part I.The Bell System Technical Journal, 37(1):185–282, 1958
1958
-
[16]
Using acoustic indices in ecology: Guidance on study design, analyses and interpretation.Methods in Ecology and Evolution, 14(9):2192–2204, 2023
Tom Bradfer-Lawrence, Camille Desjonqueres, Alice Eldridge, Alison Johnston, and Oliver Metcalf. Using acoustic indices in ecology: Guidance on study design, analyses and interpretation.Methods in Ecology and Evolution, 14(9):2192–2204, 2023
2023
-
[17]
Bar- nett, Amy Beeston, Jennifer Darby, Benedict Dell, Nick Gardner, Amandine Gasc, Becky Heath, Nia Howells, Magnus Janson, Maria-Viktoria Kyoseva, Thomas Luy- paert, Oliver C
Tom Bradfer-Lawrence, Brad Duthie, Carlos Abrahams, Maty´ aˇ s Adam, Ross J. Bar- nett, Amy Beeston, Jennifer Darby, Benedict Dell, Nick Gardner, Amandine Gasc, Becky Heath, Nia Howells, Magnus Janson, Maria-Viktoria Kyoseva, Thomas Luy- paert, Oliver C. Metcalf, Anna E. Nouse...
2025
-
[18]
Hanqun Cao, Cheng Tan, Zhangyang Gao, Yilun Xu, Guangyong Chen, Pheng-Ann Heng, and Stan Z. Li. A Survey on Generative Diffusion Models.IEEE Transactions on Knowledge and Data Engineering, 36(7):2814–2830, 2024
2024
-
[19]
Conceptual 12M: Pushing Web-Scale Image-Text Pre-Training To Recognize Long-Tail Visual Con- cepts
Soravit Changpinyo, Piyush Kumar Sharma, Nan Ding, and Radu Soricut. Conceptual 12M: Pushing Web-Scale Image-Text Pre-Training To Recognize Long-Tail Visual Con- cepts. InProceedings of the 2021 IEEE Conference on Computer Vision and Pattern Recognition, pages 3557–3567, 2021
2021
-
[20]
Phantom-Data: Towards a General Subject-Consistent Video Generation Dataset.ArXiv, 2506.18851, 2025
Zhuowei Chen, Bingchuan Li, Tianxiang Ma, Lijie Liu, Mingcong Liu, Yi Zhang, Gen Li, Xinghui Li, Siyu Zhou, Qian He, and Xinglong Wu. Phantom-Data: Towards a General Subject-Consistent Video Generation Dataset.ArXiv, 2506.18851, 2025
2025 arXiv
-
[21]
Veo, 2024
Google DeepMind. Veo, 2024. 21
2024
-
[22]
Central Limit Theorem in the Functional Approach.IEEE Transactions on Signal Processing, 61(16):4025–4037, 2013
Dominique Dehay, Jacek Leskow, and Antonio Napolitano. Central Limit Theorem in the Functional Approach.IEEE Transactions on Signal Processing, 61(16):4025–4037, 2013
2013
-
[23]
A Survey of On-Device Machine Learning: An Algorithms and Learning Theory Perspective.ACM Transactions on Internet of Things, 2(3), 2021
Sauptik Dhar, Junyao Guo, Jiayi (Jason) Liu, Samarth Tripathi, Unmesh Kurup, and Mohak Shah. A Survey of On-Device Machine Learning: An Algorithms and Learning Theory Perspective.ACM Transactions on Internet of Things, 2(3), 2021
2021
-
[24]
Jukebox: A Generative Model for Music.ArXiv, 2005.00341, 2020
Prafulla Dhariwal, Heewoo Jun, Christine Payne, Jong Wook Kim, Alec Radford, and Ilya Sutskever. Jukebox: A Generative Model for Music.ArXiv, 2005.00341, 2020
2005 arXiv
-
[25]
Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey, Abhishek Kadian, Ahmad Al- Dahle, Aiesha Letman, Akhil Mathur, Alan Schelten, Amy Yang, Angela Fan, Anirudh Goyal, Anthony Hartshorn, Aobo Yang, Archi Mitra, Archie Sravankumar, Artem Korenev, Arthur Hinsvark, Arun Rao, Aston Zh...
2024 arXiv
-
[26]
Image Generation: A Review.Neural Processing Letters, 54(5):4609–4646, 2022
Mohamed Elasri, Omar Elharrouss, Somaya Al-Maadeed, and Hamid Tairi. Image Generation: A Review.Neural Processing Letters, 54(5):4609–4646, 2022
2022
-
[27]
Scaling Rectified Flow Transformers for High-Resolution Image Synthesis
Patrick Esser, Sumith Kulal, Andreas Blattmann, Rahim Entezari, Jonas M¨ uller, Harry Saini, Yam Levi, Dominik Lorenz, Axel Sauer, Frederic Boesel, Dustin Podell, Tim Dockhorn, Zion English, Kyle Lacey, Alex Goodwin, Yannik Marek, and Robin Rombach. Scaling Rectified Flow Tran...
2024 arXiv
-
[28]
Creativity and Machine Learning: A Survey
Giorgio Franceschelli and Mirco Musolesi. Creativity and Machine Learning: A Survey. ArXiv, 2104.02726, 2022
2022 arXiv
-
[29]
The Pile: An 800GB Dataset of Diverse Text for Language Modeling.ArXiv, 2101.00027, 2020
Leo Gao, Stella Biderman, Sid Black, Laurence Golding, Travis Hoppe, Charles Foster, Jason Phang, Horace He, Anish Thite, Noa Nabeshima, Shawn Presser, and Connor Leahy. The Pile: An 800GB Dataset of Diverse Text for Language Modeling.ArXiv, 2101.00027, 2020. 24
2020 arXiv
-
[30]
Listen to Look: Action Recognition by Previewing Audio.ArXiv, 1912.04487, 2020
Ruohan Gao, Tae-Hyun Oh, Kristen Grauman, and Lorenzo Torresani. Listen to Look: Action Recognition by Previewing Audio.ArXiv, 1912.04487, 2020
1912 arXiv
-
[31]
Gemmeke, Daniel P
Jort F. Gemmeke, Daniel P. W. Ellis, Dylan Freedman, Aren Jansen, Wade Lawrence, R. Channing Moore, Manoj Plakal, and Marvin Ritter. Audio Set: An ontology and human-labeled dataset for audio events. InProceedings of the 2017 IEEE International Conference on Acoustics, Speech ...
2017
-
[32]
ImageBind: One Embedding Space To Bind Them All.ArXiv, 2305.05665, 2023
Rohit Girdhar, Alaaeldin El-Nouby, Zhuang Liu, Mannat Singh, Kalyan Vasudev Al- wala, Armand Joulin, and Ishan Misra. ImageBind: One Embedding Space To Bind Them All.ArXiv, 2305.05665, 2023
2023 arXiv
-
[33]
Kristen Grauman, Andrew Westbury, Eugene Byrne, Vincent Cartillier, Zachary Chavis, Antonino Furnari, Rohit Girdhar, Jackson Hamburger, Hao Jiang, Devansh Kukreja, Miao Liu, Xingyu Liu, Miguel Martin, Tushar Nagarajan, Ilija Radosavovic, Santhosh Kumar Ramakrishnan, Fiona Ryan...
2024
-
[34]
Mamba: Linear-Time Sequence Modeling with Selective State Spaces.ArXiv, 2312.00752, 2024
Albert Gu and Tri Dao. Mamba: Linear-Time Sequence Modeling with Selective State Spaces.ArXiv, 2312.00752, 2024
2024 arXiv
-
[35]
Temporal Alignment Networks for Long-term Video.ArXiv, 2204.02968, 2022
Tengda Han, Weidi Xie, and Andrew Zisserman. Temporal Alignment Networks for Long-term Video.ArXiv, 2204.02968, 2022
2022 arXiv
-
[36]
Rae, and Laurent Sifre
Jordan Hoffmann, Sebastian Borgeaud, Arthur Mensch, Elena Buchatskaya, Trevor Cai, Eliza Rutherford, Diego de Las Casas, Lisa Anne Hendricks, Johannes Welbl, Aidan Clark, Tom Hennigan, Eric Noland, Katie Millican, George van den Driessche, Bogdan Damoc, Aurelia Guy, Simon Osin...
2024
-
[37]
Intuitive Multilingual Audio-Visual Speech Recognition with a Single-Trained Model
Joanna Hong, Se Park, and Yong Ro. Intuitive Multilingual Audio-Visual Speech Recognition with a Single-Trained Model. InFindings of the 2023 Conference on Empirical Methods in Natural Language Processing, pages 4886–4890, 2023
2023
-
[38]
Key Frame Selection for Temporal Graph Opti- mization of Skeleton-Based Action Recognition.Applied Sciences, 14(21), 2024
Jingyi Hou, Lei Su, and Yan Zhao. Key Frame Selection for Temporal Graph Opti- mization of Skeleton-Based Action Recognition.Applied Sciences, 14(21), 2024
2024
-
[39]
NLIP: Noise-Robust Language-Image Pre-training
Runhui Huang, Yanxin Long, Jianhua Han, Hang Xu, Xiwen Liang, Chunjing Xu, and Xiaodan Liang. NLIP: Noise-Robust Language-Image Pre-training. InProceedings of the 37th AAAI Conference on Artificial Intelligence, pages 926–934, 2023
2023
-
[40]
Imagen-Team-Google, :, Jason Baldridge, Jakob Bauer, Mukul Bhutani, Nicole Brich- tova, Andrew Bunner, Kelvin Chan, Yichang Chen, Sander Dieleman, Yuqing Du, Zach Eaton-Rosen, Hongliang Fei, Nando de Freitas, Yilin Gao, Evgeny Gladchenko, Sergio G´ omez Colmenarejo, Mandy Guo,...
2024
-
[41]
Dennis L. Jackson. Revisiting Sample Size and Number of Parameter Estimates: Some Support for the N:q Hypothesis.Structural Equation Modeling: A Multidisciplinary Journal, 10(1):128–141, 2003
2003
-
[42]
LL VIP: A Visible- infrared Paired Dataset for Low-light Vision
Xinyu Jia, Chuang Zhu, Minzhen Li, Wenqi Tang, and Wenli Zhou. LL VIP: A Visible- infrared Paired Dataset for Low-light Vision. InProceedings of the 2021 IEEE Inter- national Conference on Computer Vision Workshops, pages 3489–3497, 2021
2021
-
[43]
Nicolas Jonason and Bob L. T. Sturm. TimbreCLIP: Connecting Timbre to Text and Images.ArXiv, 2211.11225, 2022
2022 arXiv
-
[44]
Noise-Aware Learning from Web-Crawled Image-Text Data for Image Captioning
Wooyoung Kang, Jonghwan Mun, Sungjun Lee, and Byungseok Roh. Noise-Aware Learning from Web-Crawled Image-Text Data for Image Captioning. InProceedings of the 2023 IEEE International Conference on Computer Vision, pages 2942–2952, 2023
2023
-
[45]
MMIS: Multimodal Dataset for Interior Scene Visual Generation and Recog- nition.ArXiv, 2407.05980, 2024
Hozaifa Kassab, Ahmed Mahmoud, Mohamed Bahaa, Ammar Mohamed, and Ali Hamdi. MMIS: Multimodal Dataset for Interior Scene Visual Generation and Recog- nition.ArXiv, 2407.05980, 2024
2024 arXiv
-
[46]
Zahra Khanjani, Gabrielle Watson, and Vandana P. Janeja. Audio deepfakes: A survey. Frontiers in Big Data, 5, 2023
2023
-
[47]
AudioCaps: Generating Captions for Audios in The Wild
Chris Dongjoo Kim, Byeongchang Kim, Hyunmin Lee, and Gunhee Kim. AudioCaps: Generating Captions for Audios in The Wild. InProceedings of the 2019 North Amer- ican Chapter of the Association for Computational Linguistics, pages 119–132, 2019
2019
-
[48]
Benchmarking Cognitive Biases in Large Language Models as Evaluators.ArXiv, 2309.17012, 2023
Ryan Koo, Minhwa Lee, Vipul Raheja, Jong Inn Park, Zae Myung Kim, and Dongyeop Kang. Benchmarking Cognitive Biases in Large Language Models as Evaluators.ArXiv, 2309.17012, 2023. 27
2023 arXiv
-
[49]
AudioGen: Textually Guided Audio Generation.ArXiv, 2209.15352, 2023
Felix Kreuk, Gabriel Synnaeve, Adam Polyak, Uriel Singer, Alexandre D´ efossez, Jade Copet, Devi Parikh, Yaniv Taigman, and Yossi Adi. AudioGen: Textually Guided Audio Generation.ArXiv, 2209.15352, 2023
2023 arXiv
-
[50]
BindDiffusion: One Diffusion Model to Bind Them All, 2024
Sea AI Lab. BindDiffusion: One Diffusion Model to Bind Them All, 2024
2024
-
[51]
FLUX, 2024
Black Forest Labs. FLUX, 2024
2024
-
[52]
Jorge E. Le´ on. A VT Multimodal Dataset, 2024
2024
-
[53]
Jorge E. Le´ on. Image-audio pairs (1 of 3), 2024
2024
-
[54]
Jorge E. Le´ on. Image-audio pairs (2 of 3), 2024
2024
-
[55]
Jorge E. Le´ on. Image-audio pairs (3 of 3), 2024
2024
-
[56]
Jorge E. Le´ on. Text-audio pairs (1 of 4), 2024
2024
-
[57]
Jorge E. Le´ on. Text-audio pairs (2 of 4), 2024
2024
-
[58]
Jorge E. Le´ on. Text-audio pairs (3 of 4), 2024
2024
-
[59]
Jorge E. Le´ on. Text-audio pairs (4 of 4), 2024
2024
-
[60]
OpenHumanVid: A Large-Scale High-Quality Dataset for Enhancing Human-Centric Video Generation
Hui Li, Mingwang Xu, Yun Zhan, Shan Mu, Jiaye Li, Kaihui Cheng, Yuxuan Chen, Tan Chen, Mao Ye, Jingdong Wang, and Siyu Zhu. OpenHumanVid: A Large-Scale High-Quality Dataset for Enhancing Human-Centric Video Generation. InProceedings of the 2025 IEEE Conference on Computer Visi...
2025
-
[61]
BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Mod- els.ArXiv, 2301.12597, 2023
Junnan Li, Dongxu Li, Silvio Savarese, and Steven Hoi. BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Mod- els.ArXiv, 2301.12597, 2023
2023 arXiv
-
[62]
BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and Gen- eration.ArXiv, 2201.12086, 2022
Junnan Li, Dongxu Li, Caiming Xiong, and Steven Hoi. BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and Gen- eration.ArXiv, 2201.12086, 2022
2022 arXiv
-
[63]
Word-Level Explanations for Analyzing Bias in Text-to-Image Models.ArXiv, 2306.05500, 2023
Alexander Lin, Lucas Monteiro Paes, Sree Harsha Tanneru, Suraj Srinivas, and Himabindu Lakkaraju. Word-Level Explanations for Analyzing Bias in Text-to-Image Models.ArXiv, 2306.05500, 2023
2023 arXiv
-
[64]
Lawrence Zitnick
Tsung-Yi Lin, Michael Maire, Serge Belongie, James Hays, Pietro Perona, Deva Ra- manan, Piotr Doll´ ar, and C. Lawrence Zitnick. Microsoft COCO: Common Objects in Context. InProceedings of the 13th European Conference on Computer Vision, pages 740–755, 2014. 28
2014
-
[65]
A Comparison Between KeyFrame Extraction Methods for Clothing Recognition, 2023
Gabriel Lindgren. A Comparison Between KeyFrame Extraction Methods for Clothing Recognition, 2023
2023
-
[66]
Plumbley
Haohe Liu, Zehua Chen, Yi Yuan, Xinhao Mei, Xubo Liu, Danilo Mandic, Wenwu Wang, and Mark D. Plumbley. AudioLDM: Text-to-Audio Generation with Latent Diffusion Models. InProceedings of the 40th International Conference on Machine Learning, pages 21450–21474, 2023
2023
-
[67]
Sora: A Review on Background, Technology, Limitations, and Opportunities of Large Vision Models.ArXiv, 2402.17177, 2024
Yixin Liu, Kai Zhang, Yuan Li, Zhiling Yan, Chujie Gao, Ruoxi Chen, Zhengqing Yuan, Yue Huang, Hanchi Sun, Jianfeng Gao, Lifang He, and Lichao Sun. Sora: A Review on Background, Technology, Limitations, and Opportunities of Large Vision Models.ArXiv, 2402.17177, 2024
2024 arXiv
-
[68]
Z. Liu, P. Luo, S. Qiu, X. Wang, and X. Tang. DeepFashion: Powering Robust Clothes Recognition and Retrieval with Rich Annotations. InProceedings of the 2016 IEEE Conference on Computer Vision and Pattern Recognition, pages 1096–1104, 2016
2016
-
[69]
BLAP: Bootstrapping Language-Audio Pre-training for Music Captioning
Nathana¨ el Perraudin Luca A Lanzend¨ orfer, Constantin Pinkl and Roger Wattenhofer. BLAP: Bootstrapping Language-Audio Pre-training for Music Captioning. InAudio Imagination: NeurIPS 2024 Workshop AI-Driven Speech, Music, and Sound Genera- tion, 2024
2024
-
[70]
Stable Diffusion Akashic Records, 2023
Maks-s. Stable Diffusion Akashic Records, 2023
2023
-
[71]
GenRL: Multimodal-foundation world models for generalization in embodied agents
Pietro Mazzaglia, Tim Verbelen, Bart Dhoedt, Aaron Courville, and Sai Rajeswar. GenRL: Multimodal-foundation world models for generalization in embodied agents. ArXiv, 2406.18043, 2024
2024 arXiv
-
[72]
Mustango: Toward Controllable Text-to-Music Generation
Jan Melechovsky, Zixun Guo, Deepanway Ghosal, Navonil Majumder, Dorien Herre- mans, and Soujanya Poria. Mustango: Toward Controllable Text-to-Music Generation. InProceedings of the 2024 North American Chapter of the Association for Computa- tional Linguistics, page 8293–8316, 2024
2024
-
[73]
Mukhamediev, Adilkhan Symagulov, Yan Kuchin, Kirill Yakunin, and Ma- rina Yelis
Ravil I. Mukhamediev, Adilkhan Symagulov, Yan Kuchin, Kirill Yakunin, and Ma- rina Yelis. From Classical Machine Learning to Deep Neural Networks: A Simplified Scientometric Review.Applied Sciences, 11(12), 2021
2021
-
[74]
DALL·E 3 System Card, 2023
OpenAI. DALL·E 3 System Card, 2023
2023
-
[75]
Video generation models as world simulators, 2024
OpenAI. Video generation models as world simulators, 2024
2024
-
[76]
Douglas O’Shaughnessy.Speech Communications: Human and Machine, volume 2. 2000
2000
-
[77]
Image-to-Image Translation: Methods and Applications.IEEE Transactions on Multimedia, 24:3859–3881, 2022
Yingxue Pang, Jianxin Lin, Tao Qin, and Zhibo Chen. Image-to-Image Translation: Methods and Applications.IEEE Transactions on Multimedia, 24:3859–3881, 2022. 29
2022
-
[78]
AudioSetZSL, 2019
Kranti Kumar Parida. AudioSetZSL, 2019
2019
-
[79]
Coordi- nated Joint Multimodal Embeddings for Generalized Audio-Visual Zero-shot Classifi- cation and Retrieval of Videos
Kranti Kumar Parida, Neeraj Matiyali, Tanaya Guha, and Gaurav Sharma. Coordi- nated Joint Multimodal Embeddings for Generalized Audio-Visual Zero-shot Classifi- cation and Retrieval of Videos. InProceedings of the 2020 IEEE Winter Conference on Applications of Computer Vision,...
2020
-
[80]
Pijanowski, Luis J
Bryan C. Pijanowski, Luis J. Villanueva-Rivera, Sarah L. Dumyahn, Almo Farina, Bernie L. Krause, Brian M. Napoletano, Stuart H. Gage, and Nadia Pieretti. Sound- scape Ecology: The Science of Sound in the Landscape.BioScience, 61(3):203–216, 2011
2011
-
[81]
Plummer, Liwei Wang, Chris M
Bryan A. Plummer, Liwei Wang, Chris M. Cervantes, Juan C. Caicedo, Julia Hocken- maier, and Svetlana Lazebnik. Flickr30k Entities: Collecting Region-to-Phrase Cor- respondences for Richer Image-to-Sentence Models. InProceedings of the 2015 IEEE International Conference on Comp...
2015
-
[82]
SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis.ArXiv, 2307.01952, 2023
Dustin Podell, Zion English, Kyle Lacey, Andreas Blattmann, Tim Dockhorn, Jonas M¨ uller, Joe Penna, and Robin Rombach. SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis.ArXiv, 2307.01952, 2023
2023 arXiv
-
[83]
Does mixing of speech signals comply with central limit theorem? International Journal of Electronics and Communications, 62(10):782–785, 2008
Rajkishore Prasad. Does mixing of speech signals comply with central limit theorem? International Journal of Electronics and Communications, 62(10):782–785, 2008
2008
-
[84]
MirrorGAN: Learning Text-To-Image Generation by Redescription
Tingting Qiao, Jing Zhang, Duanqing Xu, and Dacheng Tao. MirrorGAN: Learning Text-To-Image Generation by Redescription. InProceedings of the 2019 IEEE Con- ference on Computer Vision and Pattern Recognition, pages 1505–1514, 2019
2019
-
[85]
Learning Transferable Visual Models From Natural Lan- guage Supervision.ArXiv, 2103.00020, 2021
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sand- hini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, Gretchen Krueger, and Ilya Sutskever. Learning Transferable Visual Models From Natural Lan- guage Supervision.ArXiv, 2103.00020, 2021
2021 arXiv
-
[86]
Robust Speech Recognition via Large-Scale Weak Supervision
Alec Radford, Jong Wook Kim, Tao Xu, Greg Brockman, Christine McLeavey, and Ilya Sutskever. Robust Speech Recognition via Large-Scale Weak Supervision. In Proceedings of the 40th International Conference on Machine Learning, pages 28492– 28518, 2023
2023
-
[87]
Zero-Shot Text-to-Image Generation.ArXiv, 2102.12092, 2021
Aditya Ramesh, Mikhail Pavlov, Gabriel Goh, Scott Gray, Chelsea Voss, Alec Rad- ford, Mark Chen, and Ilya Sutskever. Zero-Shot Text-to-Image Generation.ArXiv, 2102.12092, 2021
2021 arXiv
-
[88]
Machel Reid, Nikolay Savinov, Denis Teplyashin, Dmitry Lepikhin, Timothy Lillicrap, Jean baptiste Alayrac, Radu Soricut, Angeliki Lazaridou, Orhan Firat, Julian Schrit- twieser, Ioannis Antonoglou, Rohan Anil, Sebastian Borgeaud, Andrew Dai, Katie Mil- lican, Ethan Dyer, Mia G...
2024 arXiv
-
[89]
Stable Diffusion, 2021
Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, and Bj¨ orn Om- mer. Stable Diffusion, 2021. 33
2021
-
[90]
High-Resolution Image Synthesis with Latent Diffusion Models.ArXiv, 2112.10752, 2022
Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, and Bj¨ orn Ommer. High-Resolution Image Synthesis with Latent Diffusion Models.ArXiv, 2112.10752, 2022
2022 arXiv
-
[91]
Introducing Gen-3 Alpha: A New Frontier for Video Generation, 2024
Runway. Introducing Gen-3 Alpha: A New Frontier for Video Generation, 2024
2024
-
[92]
Photorealistic Text-to-Image Diffusion Models with Deep Language Understanding
Chitwan Saharia, William Chan, Saurabh Saxena, Lala Li, Jay Whang, Emily Denton, Seyed Kamyar Seyed Ghasemipour, Burcu Karagol Ayan, S Sara Mahdavi, Raphael Gontijo-Lopes, Tim Salimans, Jonathan Ho, David J Fleet, and Mohammad Norouzi. Photorealistic Text-to-Image Diffusion Mo...
2024
-
[93]
ActionAtlas: A VideoQA Benchmark for Domain-specialized Action Recognition
Mohammadreza Salehi, Jae Sung Park, Tanush Yadav, Aditya Kusupati, Ranjay Kr- ishna, Yejin Choi, Hannaneh Hajishirzi, and Ali Farhadi. ActionAtlas: A VideoQA Benchmark for Domain-specialized Action Recognition. InProceedings of the 37th In- ternational Conference on Neural Inf...
2024
-
[94]
Comparison and Analysis of Image-to- Image Generative Adversarial Networks: A Survey.ArXiv, 2112.12625, 2022
Sagar Saxena and Mohammad Nayeem Teli. Comparison and Analysis of Image-to- Image Generative Adversarial Networks: A Survey.ArXiv, 2112.12625, 2022
2022 arXiv
-
[95]
What is noise?Geophysics, 63(4):1122–1124, 1998
John Scales and Roel Snieder. What is noise?Geophysics, 63(4):1122–1124, 1998
1998
-
[96]
LAION-5B: An open large-scale dataset for training next generation image-text models
Christoph Schuhmann, Romain Beaumont, Richard Vencu, Cade Gordon, Ross Wight- man, Mehdi Cherti, Theo Coombes, Aarush Katta, Clayton Mullis, Mitchell Worts- man, Patrick Schramowski, Srivatsa Kundurthy, Katherine Crowson, Ludwig Schmidt, Robert Kaczmarczyk, and Jenia Jitsev. L...
2022
-
[97]
LAION-400M: Open Dataset of CLIP-Filtered 400 Million Image-Text Pairs.ArXiv, 2111.02114, 2021
Christoph Schuhmann, Richard Vencu, Romain Beaumont, Robert Kaczmarczyk, Clayton Mullis, Aarush Katta, Theo Coombes, Jenia Jitsev, and Aran Komatsuzaki. LAION-400M: Open Dataset of CLIP-Filtered 400 Million Image-Text Pairs.ArXiv, 2111.02114, 2021
2021 arXiv
-
[98]
C. E. Shannon. A mathematical theory of communication.The Bell System Technical Journal, 27(3):379–423, 1948
1948
-
[99]
I Hear Your True Colors: Image Guided Audio Generation
Roy Sheffer and Yossi Adi. I Hear Your True Colors: Image Guided Audio Generation. ArXiv, 2211.03089, 2023
2023 arXiv
-
[100]
A Survey on Audio Synthesis and Audio-Visual Multimodal Processing
Zhaofeng Shi. A Survey on Audio Synthesis and Audio-Visual Multimodal Processing. ArXiv, 2108.00443, 2021
2021 arXiv
-
[101]
Audio-to-Visual Cross-Modal Generation of Birds.IEEE Access, 11:27719–27729, 2023
Joo Yong Shim, Joongheon Kim, and Jong-Kook Kim. Audio-to-Visual Cross-Modal Generation of Birds.IEEE Access, 11:27719–27729, 2023. 34
2023
-
[102]
A survey on Image Data Augmentation for Deep Learning.Journal of Big Data, 6, 2019
Connor Shorten and Taghi Khoshgoftaar. A survey on Image Data Augmentation for Deep Learning.Journal of Big Data, 6, 2019
2019
-
[103]
Outpainting Images and Videos using GANs.International Journal of Computer Trends and Tech- nology, 68(5):24–29, 2020
Shailendra Singh, Nainish Aggarwal, Udit Jain, and Hrithik Jaiswal. Outpainting Images and Videos using GANs.International Journal of Computer Trends and Tech- nology, 68(5):24–29, 2020
2020
-
[104]
Lichtenberg, and Jianxiong Xiao
Shuran Song, Samuel P. Lichtenberg, and Jianxiong Xiao. SUN RGB-D: A RGB-D scene understanding benchmark suite. InProceedings of the 2015 IEEE Conference on Computer Vision and Pattern Recognition, pages 567–576, 2015
2015
-
[105]
A survey of multimodal deep generative models
Masahiro Suzuki and Yutaka Matsuo. A survey of multimodal deep generative models. Advanced Robotics, 36(5-6):261–278, 2022
2022
-
[106]
CoDi-2: In-Context, Interleaved, and Interactive Any-to-Any Generation
Zineng Tang, Ziyi Yang, Mahmoud Khademi, Yang Liu, Chenguang Zhu, and Mohit Bansal. CoDi-2: In-Context, Interleaved, and Interactive Any-to-Any Generation. ArXiv, 2311.18775, 2023
2023 arXiv
-
[107]
Any- to-any generation via composable diffusion
Zineng Tang, Ziyi Yang, Chenguang Zhu, Michael Zeng, and Mohit Bansal. Any- to-any generation via composable diffusion. InProceedings of the 37th International Conference on Neural Information Processing Systems, pages 16083–16099, 2024
2024
-
[108]
Movie Gen: A Cast of Media Foundation Models, 2024
The Movie Gen team. Movie Gen: A Cast of Media Foundation Models, 2024
2024
-
[109]
Shamma, Gerald Friedland, Benjamin Elizalde, Karl Ni, Dou- glas Poland, Damian Borth, and Li-Jia Li
Bart Thomee, David A. Shamma, Gerald Friedland, Benjamin Elizalde, Karl Ni, Dou- glas Poland, Damian Borth, and Li-Jia Li. YFCC100M: the new data in multimedia research.Communications of the ACM, 59(2):64–73, 2016
2016
-
[110]
Fugatto 1 - Foundational Genera- tive Audio Transformer Opus 1, 2024
Rafael Valle, Rohan Badlani, Zhifeng Kong, Sang gil Lee, Arushi Goel, Sungwon Kim, Joao Felipe Santos, Shuqi Dai, Siddharth Gururani, Aya AIJa’fari, Alex Liu, Kevin Shih, Wei Ping, Huck Yang, and Bryan Catanzaro. Fugatto 1 - Foundational Genera- tive Audio Transformer Opus 1, 2024
2024
-
[111]
Learning Text-to-Video Retrieval from Image Captioning.International Journal of Computer Vision, 133:1834–1854, 2024
Lucas Ventura, Cordelia Schmid, and G¨ ul Varol. Learning Text-to-Video Retrieval from Image Captioning.International Journal of Computer Vision, 133:1834–1854, 2024
2024
-
[112]
Audio Describing Sound – What Sounds are Described and How?: Results from a Flemish case study.Journal of Audiovisual Translation, 5(2):114–133, 2022
Gert Vercauteren and Nina Reviers. Audio Describing Sound – What Sounds are Described and How?: Results from a Flemish case study.Journal of Audiovisual Translation, 5(2):114–133, 2022
2022
-
[113]
Temporally Aligned Audio for Video with Autoregression
Ilpo Viertola, Vladimir Iashin, and Esa Rahtu. Temporally Aligned Audio for Video with Autoregression. InProceedings of the 2025 IEEE International Conference on Acoustics, Speech and Signal Processing, pages 1–5, 2025. 35
2025
-
[114]
Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers.ArXiv, 2301.02111, 2023
Chengyi Wang, Sanyuan Chen, Yu Wu, Ziqiang Zhang, Long Zhou, Shujie Liu, Zhuo Chen, Yanqing Liu, Huaming Wang, Jinyu Li, Lei He, Sheng Zhao, and Furu Wei. Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers.ArXiv, 2301.02111, 2023
2023 arXiv
-
[115]
AlignNet: A Unifying Approach to Audio-Visual Alignment
Jianren Wang, Zhaoyuan Fang, and Hang Zhao. AlignNet: A Unifying Approach to Audio-Visual Alignment. InProceedings of the 2020 IEEE Winter Conference on Applications of Computer Vision, pages 3298–3306, 2020
2020
-
[116]
InternVid: A Large-scale Video-Text Dataset for Multi- modal Understanding and Generation.ArXiv, 2307.06942, 2024
Yi Wang, Yinan He, Yizhuo Li, Kunchang Li, Jiashuo Yu, Xin Ma, Xinhao Li, Guo Chen, Xinyuan Chen, Yaohui Wang, Conghui He, Ping Luo, Ziwei Liu, Yali Wang, Limin Wang, and Yu Qiao. InternVid: A Large-scale Video-Text Dataset for Multi- modal Understanding and Generation.ArXiv, ...
2024 arXiv
-
[117]
Audio- Language Datasets of Scenes and Events: A Survey.ArXiv, 2407.06947, 2024
Gijs Wijngaard, Elia Formisano, Michele Esposito, and Michel Dumontier. Audio- Language Datasets of Scenes and Events: A Survey.ArXiv, 2407.06947, 2024
2024 arXiv
-
[118]
Liu, and Hung yi Lee
Haibin Wu, Xuanjun Chen, Yi-Cheng Lin, Kai wei Chang, Ho-Lam Chung, Alexan- der H. Liu, and Hung yi Lee. Towards audio language modeling – an overview.ArXiv, 2402.13236, 2024
2024 arXiv
-
[119]
Audio-Text Models Do Not Yet Leverage Natural Language
Ho-Hsiang Wu, Oriol Nieto, Juan Pablo Bello, and Justin Salamon. Audio-Text Models Do Not Yet Leverage Natural Language. InProceedings of the 2023 IEEE International Conference on Acoustics, Speech and Signal Processing, pages 1–5, 2023
2023
-
[120]
Wav2CLIP: Learning Robust Audio Representations from Clip
Ho-Hsiang Wu, Prem Seetharaman, Kundan Kumar, and Juan Pablo Bello. Wav2CLIP: Learning Robust Audio Representations from Clip. InProceedings of the 2022 IEEE International Conference on Acoustics, Speech and Signal Processing, pages 4563–4567, 2022
2022
-
[121]
NExT-GPT: Any- to-Any Multimodal LLM.ArXiv, 2309.05519, 2024
Shengqiong Wu, Hao Fei, Leigang Qu, Wei Ji, and Tat-Seng Chua. NExT-GPT: Any- to-Any Multimodal LLM.ArXiv, 2309.05519, 2024
2024 arXiv
-
[122]
Sound- scape diversity: Evaluation indices of the sound environment in urban green spaces – Effectiveness, role, and interpretation.Ecological Indicators, 154:110725, 2023
Yi Xiang, Qi Meng, Xueyong Zhang, Mengmeng Li, Da Yang, and Yue Wu. Sound- scape diversity: Evaluation indices of the sound environment in urban green spaces – Effectiveness, role, and interpretation.Ecological Indicators, 154:110725, 2023
2023
-
[123]
Peng Xu, Xiatian Zhu, and David A. Clifton. Multimodal Learning With Transformers: A Survey.IEEE Transactions on Pattern Analysis and Machine Intelligence, pages 1–20, 2023
2023
-
[124]
Xuenan Xu, Zhiling Zhang, Zelin Zhou, Pingyue Zhang, Zeyu Xie, Mengyue Wu, and Kenny Q. Zhu. BLAT: Bootstrapping Language-Audio Pre-training based on AudioSet Tag-guided Synthetic Data. InProceedings of the 31st ACM International Conference on Multimedia, page 2756–2764, 2023. 36
2023
-
[125]
Advancing High-Resolution Video-Language Repre- sentation with Large-Scale Video Transcriptions
Hongwei Xue, Tiankai Hang, Yanhong Zeng, Yuchong Sun, Bei Liu, Huan Yang, Jianlong Fu, and Baining Guo. Advancing High-Resolution Video-Language Repre- sentation with Large-Scale Video Transcriptions. InProceedings of the 2022 IEEE International Conference on Computer Vision, ...
2022
-
[126]
UltraVideo: High-Quality UHD Video Dataset with Comprehensive Captions.ArXiv, 2506.13691, 2025
Zhucun Xue, Jiangning Zhang, Teng Hu, Haoyang He, Yinan Chen, Yuxuan Cai, Yabiao Wang, Chengjie Wang, Yong Liu, Xiangtai Li, and Dacheng Tao. UltraVideo: High-Quality UHD Video Dataset with Comprehensive Captions.ArXiv, 2506.13691, 2025
2025 arXiv
-
[127]
BiCro: Noisy Correspondence Rectification for Multi-modality Data via Bi-directional Cross-modal Similarity Consistency
Shuo Yang, Zhaopan Xu, Kai Wang, Yang You, Hongxun Yao, Tongliang Liu, and Min Xu. BiCro: Noisy Correspondence Rectification for Multi-modality Data via Bi-directional Cross-modal Similarity Consistency. InProceedings of the 2023 IEEE Conference on Computer Vision and Pattern ...
2023
-
[128]
The Dawn of LMMs: Preliminary Explorations with GPT-4V(ision)
Zhengyuan Yang, Linjie Li, Kevin Lin, Jianfeng Wang, Chung-Ching Lin, Zicheng Liu, and Lijuan Wang. The Dawn of LMMs: Preliminary Explorations with GPT-4V(ision). ArXiv, 2309.17421, 2023
2023 arXiv
-
[129]
AudioToken: Adap- tation of Text-Conditioned Diffusion Models for Audio-to-Image Generation.ArXiv, 2305.13050, 2023
Guy Yariv, Itai Gat, Lior Wolf, Yossi Adi, and Idan Schwartz. AudioToken: Adap- tation of Text-Conditioned Diffusion Models for Audio-to-Image Generation.ArXiv, 2305.13050, 2023
2023 arXiv
-
[130]
Multimodal Image Synthesis and Editing: The Generative AI Era.ArXiv, 2112.13592, 2023
Fangneng Zhan, Yingchen Yu, Rongliang Wu, Jiahui Zhang, Shijian Lu, Lingjie Liu, Adam Kortylewski, Christian Theobalt, and Eric Xing. Multimodal Image Synthesis and Editing: The Generative AI Era.ArXiv, 2112.13592, 2023
2023 arXiv
-
[131]
Text-to- image Diffusion Models in Generative AI: A Survey.ArXiv, 2303.07909, 2023
Chenshuang Zhang, Chaoning Zhang, Mengchun Zhang, and In So Kweon. Text-to- image Diffusion Models in Generative AI: A Survey.ArXiv, 2303.07909, 2023
2023 arXiv
-
[132]
A Structured Model for Action Detection
Yubo Zhang, Pavel Tokmakov, Martial Hebert, and Cordelia Schmid. A Structured Model for Action Detection. InProceedings of the 2019 IEEE Conference on Computer Vision and Pattern Recognition, pages 9967–9976, 2019
2019
-
[133]
Awesome-Video-Datasets, 2023
Yunhua Zhang, Jonatan Asketorp, and Dalu Feng. Awesome-Video-Datasets, 2023
2023
-
[134]
The Sound of Pixels
Hang Zhao, Chuang Gan, Andrew Rouditchenko, Carl Vondrick, Josh McDermott, and Antonio Torralba. The Sound of Pixels. InProceedings of the 15th European Conference on Computer Vision, page 587–604, 2018
2018
-
[135]
MUSIC Dataset from Sound of Pixels, 2018
Hang Zhao and Andrew Rouditchenko. MUSIC Dataset from Sound of Pixels, 2018
2018
-
[136]
Remote Sensing Image Generation From Audio.IEEE Geoscience and Remote Sensing Letters, 18(6):994–998, 2021
Zhiyuan Zheng, Jun Chen, Xiangtao Zheng, and Xiaoqiang Lu. Remote Sensing Image Generation From Audio.IEEE Geoscience and Remote Sensing Letters, 18(6):994–998, 2021. 37
2021
-
[137]
Deep Audio-visual Learning: A Survey.International Journal of Automation and Computing, 18:351–376, 2021
Hao Zhu, Man-Di Luo, Rui Wang, Ai-Hua Zheng, and Ran He. Deep Audio-visual Learning: A Survey.International Journal of Automation and Computing, 18:351–376, 2021
2021
-
[138]
On Some Biases Encountered in Modern Audio Quality Listening Tests - A Review.Journal of the Audio Engineering Society, 56(6):427–451, 2008
S lawomir Zieli´ nski, Francis Rumsey, and Søren Bech. On Some Biases Encountered in Modern Audio Quality Listening Tests - A Review.Journal of the Audio Engineering Society, 56(6):427–451, 2008
2008
-
[139]
Audio-to-Image Cross-Modal Generation
Maciej ˙Zelaszczyk and Jacek Ma´ ndziuk. Audio-to-Image Cross-Modal Generation. ArXiv, 2109.13354, 2021. 38
2021 arXiv
Reviewed August 5, 2026 · model on record in the stance chip above.
Discussion (0). Sign in to comment.