Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-05T05:01:36.985927Z
Paper Citation Record · LEDGER
As of 8 August 2026, this Paper Citation Record lists 100 of 139 outbound references and 1 inbound Pith citation observation for arXiv:2509.05786.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-05T05:01:36.985927Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-04T21:25:26.928879Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-08-04T21:25:28.232250Z
100 of 139 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation ec8d8242-3b6f-423d-923a-7ebb9df55834 · outbound
Effectively obtaining acoustic, visual and textual data from videos MusicLM: Generating Music From Text
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3cd07964-a3bb-4358-b095-78de3fdb81ea · outbound
Effectively obtaining acoustic, visual and textual data from videos Don’t Just Assume; Look and Answer: Overcoming Priors for Visual Question Answering
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 99c6f090-968f-400b-a719-1ec22b725617 · outbound
Effectively obtaining acoustic, visual and textual data from videos Mistral Models, 2024
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 372ebeb3-17cf-44e7-a93e-0f4d91613e8b · outbound
Effectively obtaining acoustic, visual and textual data from videos Transcripter- Generation of the transcript from audio to text using Deep Learning.International Journal of Computer Sciences and Engineering, 7(1):770–773, 2019
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fc2a9b67-4ffb-489d-8878-eb742ee8314b · outbound
Effectively obtaining acoustic, visual and textual data from videos The Claude 3 Model Family: Opus, Sonnet, Haiku, 2024
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e70fe5ab-2960-40f9-85db-628bf2a1f1c6 · outbound
Effectively obtaining acoustic, visual and textual data from videos SoundNet: Learning Sound Repre- sentations from Unlabeled Video
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation defe6cc8-1aa3-440e-b120-7fc94b5f0988 · outbound
Effectively obtaining acoustic, visual and textual data from videos Multimodal Language Analysis in the Wild: CMU-MOSEI Dataset and Interpretable Dynamic Fusion Graph
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 316c0bbb-e9d0-480e-b692-35af49aa94d9 · outbound
Effectively obtaining acoustic, visual and textual data from videos AudioSetCaps: An Enriched Audio-Caption Dataset using Auto- mated Generation Pipeline with Large Audio and Language Models
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f4562230-9781-45e6-8bc2-4ce9ac26b683 · outbound
Effectively obtaining acoustic, visual and textual data from videos Frozen in Time: A Joint Video and Image Encoder for End-to-End Retrieval
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6d290b8a-a6a8-458e-90a5-6fc08362aa87 · outbound
Effectively obtaining acoustic, visual and textual data from videos Are Mod- els Biased on Text without Gender-related Language? InProceedings of the 12th International Conference on Learning Representations, 2024
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 03119dd8-0b6d-496a-b321-f8fee36243ba · outbound
Effectively obtaining acoustic, visual and textual data from videos Ballester
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b28a154a-047f-41e8-bfa6-deb51104220b · outbound
Effectively obtaining acoustic, visual and textual data from videos Improving Image Generation with Better Captions
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2ec45e4e-f960-4ff1-b3ea-6aada080ed54 · outbound
Effectively obtaining acoustic, visual and textual data from videos RenAIssance: A Survey into AI Text-to-Image Generation in the Era of Large Model
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 02198a59-9a78-4616-90b4-f1440f833053 · outbound
Effectively obtaining acoustic, visual and textual data from videos Birhane and V
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 73c69020-6c12-4726-8872-7bfcaacb2a72 · outbound
Effectively obtaining acoustic, visual and textual data from videos Unresolved cited work
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ad5ee963-7aeb-41d9-9f81-57849769cf6a · outbound
Effectively obtaining acoustic, visual and textual data from videos Using acoustic indices in ecology: Guidance on study design, analyses and interpretation.Methods in Ecology and Evolution, 14(9):2192–2204, 2023
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bd02f02c-f045-4923-9bf2-a8c0dcc8a0b3 · outbound
Effectively obtaining acoustic, visual and textual data from videos Bar- nett, Amy Beeston, Jennifer Darby, Benedict Dell, Nick Gardner, Amandine Gasc, Becky Heath, Nia Howells, Magnus Janson, Maria-Viktoria Kyoseva, Thomas Luy- paert, Oliver C
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d509692d-bf29-4afd-9c5a-4e0081b145cd · outbound
Effectively obtaining acoustic, visual and textual data from videos Unresolved cited work
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c42c01c6-2259-40b2-8bf6-054b689513a7 · outbound
Effectively obtaining acoustic, visual and textual data from videos Conceptual 12M: Pushing Web-Scale Image-Text Pre-Training To Recognize Long-Tail Visual Con- cepts
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a1f20d76-310a-4586-a321-85045de90fa4 · outbound
Effectively obtaining acoustic, visual and textual data from videos Phantom-Data : Towards a General Subject-Consistent Video Generation Dataset
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f5cc86ea-68c2-4fc8-be17-1c99e942cdf4 · outbound
Effectively obtaining acoustic, visual and textual data from videos Veo, 2024
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6d80e0b5-c3f6-4668-94e8-56131c89436a · outbound
Effectively obtaining acoustic, visual and textual data from videos Central Limit Theorem in the Functional Approach.IEEE Transactions on Signal Processing, 61(16):4025–4037, 2013
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b8407a28-33a6-4990-8995-1dc7dbda6f68 · outbound
Effectively obtaining acoustic, visual and textual data from videos A Survey of On-Device Machine Learning: An Algorithms and Learning Theory Perspective.ACM Transactions on Internet of Things, 2(3), 2021
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8963eb35-2f12-492e-8765-595ec69aed77 · outbound
Effectively obtaining acoustic, visual and textual data from videos Jukebox: A Generative Model for Music
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1767255c-ec5e-4753-80da-4b32e29684a4 · outbound
Effectively obtaining acoustic, visual and textual data from videos The Llama 3 Herd of Models
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 31693e76-b8fc-4ca9-a1f6-a460b5f4697e · outbound
Effectively obtaining acoustic, visual and textual data from videos Image Generation: A Review.Neural Processing Letters, 54(5):4609–4646, 2022
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ca63f4b9-2338-4a79-9b2f-30cb1031323f · outbound
Effectively obtaining acoustic, visual and textual data from videos Scaling Rectified Flow Transformers for High-Resolution Image Synthesis
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 38f1b52f-fbff-4cd1-958c-05edc33f38c5 · outbound
Effectively obtaining acoustic, visual and textual data from videos Creativity and Machine Learning: A Survey
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b89344f3-fc01-4770-b056-cf9382f410a9 · outbound
Effectively obtaining acoustic, visual and textual data from videos The Pile: An 800GB Dataset of Diverse Text for Language Modeling
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2c00d998-a181-45f2-900c-f5f5773eb485 · outbound
Effectively obtaining acoustic, visual and textual data from videos Listen to Look: Action Recognition by Previewing Audio
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 6b0bc007-5648-4acb-a8c1-f6965de53a81 · outbound
Effectively obtaining acoustic, visual and textual data from videos Gemmeke, Daniel P
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cc9105c9-c3f9-462e-85db-995cc7d6b868 · outbound
Effectively obtaining acoustic, visual and textual data from videos ImageBind: One Embedding Space To Bind Them All
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1d6d9822-6e67-4852-adcb-f8fbcddbb45e · outbound
Effectively obtaining acoustic, visual and textual data from videos Unresolved cited work
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 324a05e5-c1e6-4f91-a830-2717aba335bf · outbound
Effectively obtaining acoustic, visual and textual data from videos Mamba: Linear-Time Sequence Modeling with Selective State Spaces
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3ebed149-c399-4210-83ca-584ff496e8f1 · outbound
Effectively obtaining acoustic, visual and textual data from videos Temporal Alignment Networks for Long-term Video
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation f68d5808-7bb5-4605-a3ca-809bfd1eb156 · outbound
Effectively obtaining acoustic, visual and textual data from videos Rae, and Laurent Sifre
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c08dc387-7c57-484b-95b9-115f89fbfe98 · outbound
Effectively obtaining acoustic, visual and textual data from videos Intuitive Multilingual Audio-Visual Speech Recognition with a Single-Trained Model
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1a400fbb-7282-4dcf-9cc2-f6a894c37edf · outbound
Effectively obtaining acoustic, visual and textual data from videos Key Frame Selection for Temporal Graph Opti- mization of Skeleton-Based Action Recognition.Applied Sciences, 14(21), 2024
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e2619ad0-e7db-43b4-b75d-10b811b3eb7e · outbound
Effectively obtaining acoustic, visual and textual data from videos NLIP: Noise-Robust Language-Image Pre-training
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8030a54d-4642-4e07-ac9d-8312ffd4ec29 · outbound
Effectively obtaining acoustic, visual and textual data from videos Unresolved cited work
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d5b226b3-6510-48a0-9811-c9febbe51052 · outbound
Effectively obtaining acoustic, visual and textual data from videos Unresolved cited work
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 947368cd-25dd-45e2-8879-1fd29ceb871a · outbound
Effectively obtaining acoustic, visual and textual data from videos LL VIP: A Visible- infrared Paired Dataset for Low-light Vision
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 93caa914-ba11-49d9-8659-e4964ef9e460 · outbound
Effectively obtaining acoustic, visual and textual data from videos TimbreCLIP: Connecting Timbre to Text and Images
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1f7baff1-df83-4c62-853a-066e913ca8a9 · outbound
Effectively obtaining acoustic, visual and textual data from videos Noise-Aware Learning from Web-Crawled Image-Text Data for Image Captioning
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 78a115d0-1f88-48b5-ae44-df72cada4263 · outbound
Effectively obtaining acoustic, visual and textual data from videos MMIS: Multimodal Dataset for Interior Scene Visual Generation and Recognition
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 5a4026dd-2113-40a0-9e48-6acf31061119 · outbound
Effectively obtaining acoustic, visual and textual data from videos Unresolved cited work
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bc087ea3-9c7c-48f1-bec4-ac70b048557b · outbound
Effectively obtaining acoustic, visual and textual data from videos AudioCaps: Generating Captions for Audios in The Wild
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4e1e2dec-7704-4b8c-935e-200e8f2c817e · outbound
Effectively obtaining acoustic, visual and textual data from videos Benchmarking Cognitive Biases in Large Language Models as Evaluators
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a1260ad0-06c3-4bd2-87ee-c4e0c0a3b362 · outbound
Effectively obtaining acoustic, visual and textual data from videos AudioGen: Textually Guided Audio Generation
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 40ea535a-3d4a-45f8-96d2-14b38f8a8f55 · outbound
Effectively obtaining acoustic, visual and textual data from videos BindDiffusion: One Diffusion Model to Bind Them All, 2024
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e99ea385-2b6e-4480-b21e-ab166308de09 · outbound
Effectively obtaining acoustic, visual and textual data from videos FLUX, 2024
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0faaade5-8f3f-4af3-a586-0d186385b1ab · outbound
Effectively obtaining acoustic, visual and textual data from videos Unresolved cited work
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a61c7e02-bda4-4af6-9a4d-542b7246ea2c · outbound
Effectively obtaining acoustic, visual and textual data from videos Unresolved cited work
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation feb64b77-e522-4dca-b0aa-798fe3fb0323 · outbound
Effectively obtaining acoustic, visual and textual data from videos Unresolved cited work
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 23acdd7b-6340-44a0-a435-65b815bff5fc · outbound
Effectively obtaining acoustic, visual and textual data from videos Unresolved cited work
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 566fad45-6b69-461c-9233-29cfb1f62629 · outbound
Effectively obtaining acoustic, visual and textual data from videos Unresolved cited work
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 61148e1b-803c-4613-b3ff-dd1dd8afa080 · outbound
Effectively obtaining acoustic, visual and textual data from videos Unresolved cited work
Reference 57
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5074745a-3fc1-4556-95af-4a061d8d6e5e · outbound
Effectively obtaining acoustic, visual and textual data from videos Unresolved cited work
Reference 58
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dd476d19-c547-42a2-b863-03bd1bda0ba6 · outbound
Effectively obtaining acoustic, visual and textual data from videos Unresolved cited work
Reference 59
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 77b59946-7569-4cfe-8d9f-76164e9c046d · outbound
Effectively obtaining acoustic, visual and textual data from videos OpenHumanVid: A Large-Scale High-Quality Dataset for Enhancing Human-Centric Video Generation
Reference 60
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fc848145-e1df-42a1-9291-03dd77cb3ef1 · outbound
Effectively obtaining acoustic, visual and textual data from videos BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models
Reference 61
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 655e029f-4504-4828-b44b-460328908220 · outbound
Effectively obtaining acoustic, visual and textual data from videos BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and Generation
Reference 62
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 02170d03-b9a9-48b0-9946-3de063dc4f0b · outbound
Effectively obtaining acoustic, visual and textual data from videos Word-Level Explanations for Analyzing Bias in Text-to-Image Models
Reference 63
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bac5b7e1-71c3-425a-ad31-eb94a85b5459 · outbound
Effectively obtaining acoustic, visual and textual data from videos Lawrence Zitnick
Reference 64
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 724e1ee9-d848-4c4b-99ea-2e5f3deec07c · outbound
Effectively obtaining acoustic, visual and textual data from videos A Comparison Between KeyFrame Extraction Methods for Clothing Recognition, 2023
Reference 65
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3644ee48-de0c-47c8-bb57-9eed6a07f0d3 · outbound
Effectively obtaining acoustic, visual and textual data from videos Plumbley
Reference 66
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 734295a9-b47e-49ef-a458-dffdfe7d8398 · outbound
Effectively obtaining acoustic, visual and textual data from videos Sora: A Review on Background, Technology, Limitations, and Opportunities of Large Vision Models
Reference 67
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ba28e53f-7947-48ae-8690-28f0e7e5beb0 · outbound
Effectively obtaining acoustic, visual and textual data from videos Unresolved cited work
Reference 68
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 94c9feb5-863e-4d14-ba60-ec85cecbc01e · outbound
Effectively obtaining acoustic, visual and textual data from videos BLAP: Bootstrapping Language-Audio Pre-training for Music Captioning
Reference 69
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e4fd3a40-1031-4c09-8502-3071e09ade49 · outbound
Effectively obtaining acoustic, visual and textual data from videos Stable Diffusion Akashic Records, 2023
Reference 70
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9af69ee2-41fc-4bb7-9a40-bea10e6dbff3 · outbound
Effectively obtaining acoustic, visual and textual data from videos GenRL: Multimodal-foundation world models for generalization in embodied agents
Reference 71
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dce89b1f-caf0-430f-a02a-6568fa5b4c7e · outbound
Effectively obtaining acoustic, visual and textual data from videos Mustango: Toward Controllable Text-to-Music Generation
Reference 72
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e2bddc2e-62e9-4019-8434-618c75b22c38 · outbound
Effectively obtaining acoustic, visual and textual data from videos Mukhamediev, Adilkhan Symagulov, Yan Kuchin, Kirill Yakunin, and Ma- rina Yelis
Reference 73
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 95916d8e-de34-48b0-ba78-b829705542a5 · outbound
Effectively obtaining acoustic, visual and textual data from videos DALL·E 3 System Card, 2023
Reference 74
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f1955f34-dc82-4f77-897c-572158524ffe · outbound
Effectively obtaining acoustic, visual and textual data from videos Video generation models as world simulators, 2024
Reference 75
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 232e3c2f-2a47-448c-ae2f-a304e18c9c82 · outbound
Effectively obtaining acoustic, visual and textual data from videos Unresolved cited work
Reference 76
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 42c840a9-50bc-43b6-9010-7f8bb454a298 · outbound
Effectively obtaining acoustic, visual and textual data from videos Image-to-Image Translation: Methods and Applications.IEEE Transactions on Multimedia, 24:3859–3881, 2022
Reference 77
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 071a64e6-495c-4522-bfe9-fbbd871374f7 · outbound
Effectively obtaining acoustic, visual and textual data from videos AudioSetZSL, 2019
Reference 78
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2223b9a0-d786-4723-b199-03d5d455ed09 · outbound
Effectively obtaining acoustic, visual and textual data from videos Coordi- nated Joint Multimodal Embeddings for Generalized Audio-Visual Zero-shot Classifi- cation and Retrieval of Videos
Reference 79
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4ccd66b9-3ce4-4bb6-8ba1-378cccafe673 · outbound
Effectively obtaining acoustic, visual and textual data from videos Pijanowski, Luis J
Reference 80
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cbee204a-803b-4c13-bc2c-e870549cd3b2 · outbound
Effectively obtaining acoustic, visual and textual data from videos Plummer, Liwei Wang, Chris M
Reference 81
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation df676aed-bb72-4ac0-b295-29901731b91d · outbound
Effectively obtaining acoustic, visual and textual data from videos SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis
Reference 82
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d13fe9b0-f7e7-44bd-b383-7d52c3a1a9aa · outbound
Effectively obtaining acoustic, visual and textual data from videos Does mixing of speech signals comply with central limit theorem? International Journal of Electronics and Communications, 62(10):782–785, 2008
Reference 83
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 27f08d89-5358-4c74-84b1-f2cb26e6d60a · outbound
Effectively obtaining acoustic, visual and textual data from videos MirrorGAN: Learning Text-To-Image Generation by Redescription
Reference 84
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e3c07864-0d86-489c-95ab-81798b370b51 · outbound
Effectively obtaining acoustic, visual and textual data from videos Learning Transferable Visual Models From Natural Language Supervision
Reference 85
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation aedb676f-01de-41a3-996e-a5061d3e48ee · outbound
Effectively obtaining acoustic, visual and textual data from videos Robust Speech Recognition via Large-Scale Weak Supervision
Reference 86
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 87d9060f-42ac-4c74-b9f8-de95ef1ef747 · outbound
Effectively obtaining acoustic, visual and textual data from videos Zero-Shot Text-to-Image Generation
Reference 87
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f0de62db-f4ac-4b36-87d6-36ac83a0c53d · outbound
Effectively obtaining acoustic, visual and textual data from videos Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context
Reference 88
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation becfe866-8825-413b-8892-eb51594631fe · outbound
Effectively obtaining acoustic, visual and textual data from videos Stable Diffusion, 2021
Reference 89
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 0465db97-e607-4ada-a9ba-331eec2aa695 · outbound
Effectively obtaining acoustic, visual and textual data from videos High-Resolution Image Synthesis with Latent Diffusion Models
Reference 90
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8e505e95-660f-4781-bda9-a9e2b108ebb9 · outbound
Effectively obtaining acoustic, visual and textual data from videos Introducing Gen-3 Alpha: A New Frontier for Video Generation, 2024
Reference 91
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 74544f37-caf6-4c0a-ba80-96e6065ba4e1 · outbound
Effectively obtaining acoustic, visual and textual data from videos Photorealistic Text-to-Image Diffusion Models with Deep Language Understanding
Reference 92
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3fd18f94-87ce-447f-86a3-1f94f9743bd8 · outbound
Effectively obtaining acoustic, visual and textual data from videos ActionAtlas: A VideoQA Benchmark for Domain-specialized Action Recognition
Reference 93
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 6c2abd28-5c1e-43b3-abfb-6a13cf33016d · outbound
Effectively obtaining acoustic, visual and textual data from videos Comparison and Analysis of Image-to-Image Generative Adversarial Networks: A Survey
Reference 94
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1e99fdc3-62d6-4ba8-beb5-859e76015007 · outbound
Effectively obtaining acoustic, visual and textual data from videos What is noise?Geophysics, 63(4):1122–1124, 1998
Reference 95
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation f86c7b08-e085-4929-a849-b350da166928 · outbound
Effectively obtaining acoustic, visual and textual data from videos LAION-5B: An open large-scale dataset for training next generation image-text models
Reference 96
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 1141ad1b-87bf-4cf4-8205-39e917d797f9 · outbound
Effectively obtaining acoustic, visual and textual data from videos LAION-400M: Open Dataset of CLIP-Filtered 400 Million Image-Text Pairs
Reference 97
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 181fc8f0-d346-40ef-867e-3e508f577b82 · outbound
Effectively obtaining acoustic, visual and textual data from videos Unresolved cited work
Reference 98
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7bae9aa2-c149-499e-a4ec-8a463198ecb6 · outbound
Effectively obtaining acoustic, visual and textual data from videos I Hear Your True Colors: Image Guided Audio Generation
Reference 99
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b956306a-aaff-4ab1-bd9c-34721c77bd74 · outbound
Effectively obtaining acoustic, visual and textual data from videos A Survey on Audio Synthesis and Audio-Visual Multimodal Processing
Reference 100
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7f772987-5f54-46cf-9866-ec6efe4b5788 · inbound
Testing chatbots on the creation of encoders for audio conditioned image generation Effectively obtaining acoustic, visual and textual data from videos
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.