Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-04T12:39:00.193571Z
Paper Citation Record · LEDGER
As of 9 August 2026, this Paper Citation Record lists 34 of 34 outbound references and 4 inbound Pith citation observations for arXiv:2510.03117.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-04T12:39:00.193571Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-06-28T15:37:32.767405Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-07-01T22:16:16.088895Z
34 of 34 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation f6d1085d-4c0f-4aaa-92ba-2decbfb146aa · outbound
Taming Text-to-Sounding Video Generation via Advanced Modality Condition and Interaction Qwen2.5-VL Technical Report
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cff80d68-0d8b-4dae-bd22-293af6b01781 · outbound
Taming Text-to-Sounding Video Generation via Advanced Modality Condition and Interaction Clap learning audio concepts from natural language supervision
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 914d8b6e-014e-40d0-84c3-b0e5e2701f6f · outbound
Taming Text-to-Sounding Video Generation via Advanced Modality Condition and Interaction Stable audio open
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 26c57423-8f68-49ee-9012-7e0b38030f66 · outbound
Taming Text-to-Sounding Video Generation via Advanced Modality Condition and Interaction ETVA: Evaluation of Text-to-Video Alignment via Fine-grained Question Generation and Answering
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 295887da-3bab-4ee3-bbf4-dd8932dc092b · outbound
Taming Text-to-Sounding Video Generation via Advanced Modality Condition and Interaction A Simple but Strong Baseline for Sounding Video Generation: Effective Adaptation of Audio and Video Diffusion Models for Joint Generation
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4dca9b91-a497-4647-95c1-47c63dbb5973 · outbound
Taming Text-to-Sounding Video Generation via Advanced Modality Condition and Interaction Auto-Encoding Variational Bayes
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c92327c4-f020-4527-9237-4970f7d13999 · outbound
Taming Text-to-Sounding Video Generation via Advanced Modality Condition and Interaction Understanding Sounds, Missing the Questions: The Challenge of Object Hallucination in Large Audio-Language Models
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 34bee314-c44f-4b11-9eb2-b14abdf8bebe · outbound
Taming Text-to-Sounding Video Generation via Advanced Modality Condition and Interaction Sound-guided semantic video generation
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c586ab46-a947-48f8-b992-f7cbc7d348af · outbound
Taming Text-to-Sounding Video Generation via Advanced Modality Condition and Interaction Open-Sora Plan: Open-Source Large Video Generation Model
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 007728a7-707a-4f8d-9c87-3944ee18eaef · outbound
Taming Text-to-Sounding Video Generation via Advanced Modality Condition and Interaction Flow Matching for Generative Modeling
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cf197551-189e-4413-bfbb-75800a2a5d56 · outbound
Taming Text-to-Sounding Video Generation via Advanced Modality Condition and Interaction AudioLDM: Text-to-Audio Generation with Latent Diffusion Models
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 42b6baa9-358b-4bba-97f5-d90724595e5d · outbound
Taming Text-to-Sounding Video Generation via Advanced Modality Condition and Interaction SyncFlow: Toward Temporally Aligned Joint Audio-Video Generation from Text
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation aa6800e8-532a-4740-b0e4-4bc2e4687a68 · outbound
Taming Text-to-Sounding Video Generation via Advanced Modality Condition and Interaction On the Audio Hallucinations in Large Audio-Video Language Models
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 55358733-e763-4e3a-98b1-45ca7a1da5a8 · outbound
Taming Text-to-Sounding Video Generation via Advanced Modality Condition and Interaction Scalable Diffusion Models with Transformers
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 596cd1fa-c47f-4aef-92c1-d4af1235c8dd · outbound
Taming Text-to-Sounding Video Generation via Advanced Modality Condition and Interaction Progressive Distillation for Fast Sampling of Diffusion Models
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 89fd54c6-55f7-4915-982e-afac26dd7711 · outbound
Taming Text-to-Sounding Video Generation via Advanced Modality Condition and Interaction AVHBench: A Cross-Modal Hallucination Benchmark for Audio-Visual Large Language Models
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 88a031f4-e492-4d90-9bbb-c414ab844e1b · outbound
Taming Text-to-Sounding Video Generation via Advanced Modality Condition and Interaction Atom of thoughts for markov llm test-time scaling.arXiv preprint arXiv:2502.12018,
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2c3c091a-23c9-4a18-8790-7f1f7a60d5a1 · outbound
Taming Text-to-Sounding Video Generation via Advanced Modality Condition and Interaction Towards Accurate Generative Models of Video: A New Metric & Challenges
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4125dad3-9ae4-4b57-a919-f3c7a845552b · outbound
Taming Text-to-Sounding Video Generation via Advanced Modality Condition and Interaction Unresolved cited work
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d0ec8403-afbd-4bd5-8376-38d561ef2c3f · outbound
Taming Text-to-Sounding Video Generation via Advanced Modality Condition and Interaction HunyuanVideo: A Systematic Framework For Large Video Generative Models
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6af38f93-5e40-4170-9be7-204739b9be4c · outbound
Taming Text-to-Sounding Video Generation via Advanced Modality Condition and Interaction Qwen-Image Technical Report
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f022d557-bf23-4f7e-953e-ee71902b74ad · outbound
Taming Text-to-Sounding Video Generation via Advanced Modality Condition and Interaction Qwen2.5-Omni Technical Report
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2ebabfad-8135-401c-ad68-f327bf9b834e · outbound
Taming Text-to-Sounding Video Generation via Advanced Modality Condition and Interaction Qwen2.5 Technical Report
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fe916510-e314-4523-bf4a-49e286c8d953 · outbound
Taming Text-to-Sounding Video Generation via Advanced Modality Condition and Interaction UniForm: A Unified Multi-Task Diffusion Transformer for Audio-Video Generation
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 04060fd3-7145-471d-96fb-d3432a5d9c37 · outbound
Taming Text-to-Sounding Video Generation via Advanced Modality Condition and Interaction Open-Sora: Democratizing Efficient Video Production for All
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6b6a6274-63ef-4a1d-bc98-7fe359387da2 · outbound
Taming Text-to-Sounding Video Generation via Advanced Modality Condition and Interaction This technique steers the generation pro- cess towards a desired conditionc(e.g., a text prompt) without needing an external classifier
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d3105885-199f-46f0-ad1f-7ea8fa4ef1b7 · outbound
Taming Text-to-Sounding Video Generation via Advanced Modality Condition and Interaction 3D U-Net: Learning Dense Volumetric Segmentation from Sparse Annotation
Reference 2016
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d25f7a8e-4b31-46dd-bab6-4a479a67d7ec · outbound
Taming Text-to-Sounding Video Generation via Advanced Modality Condition and Interaction Wan: Open and Advanced Large-Scale Video Generative Models
Reference 2018
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b2e16779-9699-405d-8958-8f66d33a3c9f · outbound
Taming Text-to-Sounding Video Generation via Advanced Modality Condition and Interaction Denoising Diffusion Probabilistic Models
Reference 2020
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 23e273f7-3da5-4cf4-9ee3-4e4cefa5f5fa · outbound
Taming Text-to-Sounding Video Generation via Advanced Modality Condition and Interaction Qwen2-Audio Technical Report
Reference 2021
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e8144d74-5a43-4bb6-85dd-216833630784 · outbound
Taming Text-to-Sounding Video Generation via Advanced Modality Condition and Interaction Fr\'echet Audio Distance: A Metric for Evaluating Music Enhancement Algorithms
Reference 2022
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a00d6439-4197-49e2-a6c5-bfd1049acff1 · outbound
Taming Text-to-Sounding Video Generation via Advanced Modality Condition and Interaction Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets
Reference 2023
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 11f089b3-86bc-4e12-8093-03ad70f608d8 · outbound
Taming Text-to-Sounding Video Generation via Advanced Modality Condition and Interaction TEMPO: Prompt-based Generative Pre-trained Transformer for Time Series Forecasting
Reference 2024
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4fc2be78-903a-4988-b3c2-2a40cfc66194 · outbound
Taming Text-to-Sounding Video Generation via Advanced Modality Condition and Interaction Classifier-Free Diffusion Guidance
Reference 2025
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ba269173-b442-4f69-8027-50f31d2a4c67 · inbound
LTX-2: Efficient Joint Audio-Visual Foundation Model Taming Text-to-Sounding Video Generation via Advanced Modality Condition and Interaction
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 82aa4fc2-5520-4e25-8caa-efde5e885ae3 · inbound
Script-a-Video: Deep Structured Audio-visual Captions via Factorized Streams and Relational Grounding Taming Text-to-Sounding Video Generation via Advanced Modality Condition and Interaction
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 8d553e74-c230-462f-91a2-63a9399b64b8 · inbound
SyncDPO: Enhancing Temporal Synchronization in Video-Audio Joint Generation via Preference Learning Taming Text-to-Sounding Video Generation via Advanced Modality Condition and Interaction
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 4f678117-5b37-4ccd-8b31-428359713c1e · inbound
Planar Symmetric Pattern Generation Taming Text-to-Sounding Video Generation via Advanced Modality Condition and Interaction
Reference 73
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.