Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T19:40:06.423278Z
Paper Citation Record · LEDGER
As of 8 August 2026, this Paper Citation Record lists 54 of 54 outbound references and 1 inbound Pith citation observation for arXiv:2507.04955.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T19:40:06.423278Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-06T19:40:02.091028Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-08-06T19:40:07.168991Z
54 of 54 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 9e40e093-95c3-427f-b276-9e5220945e43 · outbound
EXPOTION: Facial Expression and Motion Control for Multimodal Music Generation EXPOTION: Facial Expression and Motion Control for Multimodal Music Generation
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation a5161bbe-ecc6-451f-9590-4f7a5a01a468 · outbound
EXPOTION: Facial Expression and Motion Control for Multimodal Music Generation Visual and Motion-Based Control Early interactive sys- tems mapped facial or bodily features directly to sound
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation f50d091c-f22a-40d2-8453-b09dcb9ae82a · outbound
EXPOTION: Facial Expression and Motion Control for Multimodal Music Generation We froze the parame- ters of the vanilla Musicgen during training to preserve its text understanding ability
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation c1dde247-c637-4b46-943d-0d3213644c8a · outbound
EXPOTION: Facial Expression and Motion Control for Multimodal Music Generation We recruited volunteers to record their facial expressions and upper body movements while listening to 30-second audio clips
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 28781a49-92c9-4ce2-9d9a-24267fca8f1c · outbound
EXPOTION: Facial Expression and Motion Control for Multimodal Music Generation Unresolved cited work
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 33c5a219-3486-401f-9e06-230fdbce7c2e · outbound
EXPOTION: Facial Expression and Motion Control for Multimodal Music Generation Unresolved cited work
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 5411faeb-ed22-49a3-b1a1-44ddb2381623 · outbound
EXPOTION: Facial Expression and Motion Control for Multimodal Music Generation Taming Multimodal Joint Training for High-Quality Video-to-Audio Synthe- sis,
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 9edde2b6-174f-4fb3-9ae6-bb3ffff98c6f · outbound
EXPOTION: Facial Expression and Motion Control for Multimodal Music Generation Diff-Foley: Syn- chronized Video-to-Audio Synthesis with Latent Dif- fusion Models,
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 280c36b4-c927-457d-bcaa-f567ebd5a954 · outbound
EXPOTION: Facial Expression and Motion Control for Multimodal Music Generation Temporally Aligned Audio for Video with Autoregression,
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 6607a917-256c-4f2f-bd66-2a9b9f04c558 · outbound
EXPOTION: Facial Expression and Motion Control for Multimodal Music Generation FoleyCrafter: Bring Silent Videos to Life with Lifelike and Synchronized Sounds
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7ca01989-1023-4bec-a53f-64ee01474ca4 · outbound
EXPOTION: Facial Expression and Motion Control for Multimodal Music Generation Frieren: Efficient Video-to-Audio Generation Network with Rectified Flow Matching,
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 6d1a5a91-9a00-4e55-92d7-7bf31a0b806f · outbound
EXPOTION: Facial Expression and Motion Control for Multimodal Music Generation MM-LDM: Multi-Modal Latent Diffusion Model for Sounding Video Generation,
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 3591a8b3-2211-42de-a461-5c3226339e1f · outbound
EXPOTION: Facial Expression and Motion Control for Multimodal Music Generation Visual Echoes: A Simple Unified Transformer for Audio-Visual Generation
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 347c438d-e88c-4438-a8f0-87b421fa314a · outbound
EXPOTION: Facial Expression and Motion Control for Multimodal Music Generation Video-Foley: Two-Stage Video-To-Sound Generation via Temporal Event Condition For Foley Sound,
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 98613fea-9f29-40f2-a7ed-af9861c212ed · outbound
EXPOTION: Facial Expression and Motion Control for Multimodal Music Generation InstaGen: Enhancing Object Detection by Training on Synthetic Dataset
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 1914b229-406f-4b65-95f1-44bd64a2e926 · outbound
EXPOTION: Facial Expression and Motion Control for Multimodal Music Generation Orbital angular momentum of Bloch electrons: equilibrium formulation, magneto-electric phenomena, and the orbital Hall effect
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 3268cad6-05c1-4629-991e-60b6452c29b3 · outbound
EXPOTION: Facial Expression and Motion Control for Multimodal Music Generation Conditional Generation of Audio from Video via Fo- ley Analogies,
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 570ab3ab-09c6-4f44-a80b-eb1002082a44 · outbound
EXPOTION: Facial Expression and Motion Control for Multimodal Music Generation MuVi: Video-to-Music Generation with Semantic Alignment and Rhythmic Synchronization
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5c4325b2-3013-4bf5-94e1-01675403956d · outbound
EXPOTION: Facial Expression and Motion Control for Multimodal Music Generation VMAs: Video-to-Music Generation via Semantic Alignment in Web Music Videos,
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation ace31055-2ce3-4c57-9ada-33435a441f6d · outbound
EXPOTION: Facial Expression and Motion Control for Multimodal Music Generation Video2Music: Suitable Music Generation from Videos using an Affec- tive Multimodal Transformer model,
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 272d2f30-f92d-4de2-88de-12a9bd7183b7 · outbound
EXPOTION: Facial Expression and Motion Control for Multimodal Music Generation V2Meow: Meowing to the Visual Beat via Video-to-Music Generation,
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation df50cb85-48d6-4390-8928-3e959bdb753e · outbound
EXPOTION: Facial Expression and Motion Control for Multimodal Music Generation DanceCom- poser: Dance-to-Music Generation Using a Progressive Conditional Music Generator,
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 7e4af507-c12d-48d4-ad6b-2f3f28224bbe · outbound
EXPOTION: Facial Expression and Motion Control for Multimodal Music Generation Discrete Contrastive Diffusion for Cross- Modal Music and Image Generation,
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 7c3a79c2-7f4f-4cd3-a3e4-95985327a9f3 · outbound
EXPOTION: Facial Expression and Motion Control for Multimodal Music Generation Long- Term Rhythmic Video Soundtracker,
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 20b4b726-aaae-4943-aaff-ce82af8fc7c4 · outbound
EXPOTION: Facial Expression and Motion Control for Multimodal Music Generation Simple and control- lable music generation,
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 8f0bba4b-2550-4491-86ea-2112d40c1c49 · outbound
EXPOTION: Facial Expression and Motion Control for Multimodal Music Generation Content-based Controls For Music Large Language Modeling
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 867e4e4d-3533-4f57-9782-533fc52b2236 · outbound
EXPOTION: Facial Expression and Motion Control for Multimodal Music Generation BiMediX: Bilingual Medical Mixture of Experts LLM
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 95569952-453b-4786-a063-b217798b55b4 · outbound
EXPOTION: Facial Expression and Motion Control for Multimodal Music Generation Audioclip: Extending clip to image, text and audio,
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 660763c1-7f85-49a8-abe1-d7d88733bbf8 · outbound
EXPOTION: Facial Expression and Motion Control for Multimodal Music Generation Exploring the limits of transfer learning with a unified text-to-text transformer,
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1a77d445-28de-4259-85b5-7470c0e565af · outbound
EXPOTION: Facial Expression and Motion Control for Multimodal Music Generation Vidmuse: A simple video-to- music generation framework with long-short-term mod- eling,
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation b3655a34-319f-49d9-8e80-239d6fc3e059 · outbound
EXPOTION: Facial Expression and Motion Control for Multimodal Music Generation Sonify Your Face: Facial Expressions for Sound Generation,
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 213f40a7-211f-4da6-aeae-8ee1e726f045 · outbound
EXPOTION: Facial Expression and Motion Control for Multimodal Music Generation Movement to emotions to music: using whole body emotional expression as an interaction for electronic music genera- tion,
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 5a30467f-308d-41b2-b8e5-4398e4ffedfb · outbound
EXPOTION: Facial Expression and Motion Control for Multimodal Music Generation D2MNet for music generation jointly driven by facial expressions and dance movements,
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 8beac812-bcff-45fc-9d2a-dfb7b735dc7c · outbound
EXPOTION: Facial Expression and Motion Control for Multimodal Music Generation Deep- Tunes: Music Generation based on Facial Emotions using Deep Learning,
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 95e5e6ef-e7cd-4911-b0e4-dd7226c2d195 · outbound
EXPOTION: Facial Expression and Motion Control for Multimodal Music Generation A Contin- uous Emotional Music Generation System Based on Facial Expressions,
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 290f3bdf-e9f3-492b-8641-be59b2b62b56 · outbound
EXPOTION: Facial Expression and Motion Control for Multimodal Music Generation Musiclm: Generating music from text,
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 437cbddc-94b7-4e31-b553-70831a24277a · outbound
EXPOTION: Facial Expression and Motion Control for Multimodal Music Generation Audio set: An on- tology and human-labeled dataset for audio events,
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 8d4a48bf-9cba-4a3e-a001-b3488749bedb · outbound
EXPOTION: Facial Expression and Motion Control for Multimodal Music Generation Vg- gsound: A large-scale audio-visual dataset,
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 22542eab-239f-45f9-a1d1-3db28816f211 · outbound
EXPOTION: Facial Expression and Motion Control for Multimodal Music Generation Arrange, inpaint, and refine: Steerable long-term music audio generation and editing via content-based controls,
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation c29f1b56-bd19-4960-bb1b-e28fac85f60b · outbound
EXPOTION: Facial Expression and Motion Control for Multimodal Music Generation Marlin: Masked autoencoder for facial video representation learning,
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 83659852-de37-4dcc-a2fc-0941b0e9eaea · outbound
EXPOTION: Facial Expression and Motion Control for Multimodal Music Generation Synch- former: Efficient synchronization from sparse cues,
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 8ee1745e-fb97-48c2-8d7c-8b64bbcdc768 · outbound
EXPOTION: Facial Expression and Motion Control for Multimodal Music Generation Raft: Recurrent all-pairs field transforms for optical flow,
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 30df38f1-ba3e-4b42-9166-3b4bc2c8f2e6 · outbound
EXPOTION: Facial Expression and Motion Control for Multimodal Music Generation Salmonn: Towards generic hear- ing abilities for large language models,
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation a906a825-ceb7-4767-a58c-469be428b1c0 · outbound
EXPOTION: Facial Expression and Motion Control for Multimodal Music Generation High Fidelity Neural Audio Compression
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3570c2dd-fc12-4839-94ba-a278eac465da · outbound
EXPOTION: Facial Expression and Motion Control for Multimodal Music Generation Video2music: Suitable music generation from videos using an affective multimodal transformer model,
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 4d498afb-aabf-454d-b1f7-7cb6a779273d · outbound
EXPOTION: Facial Expression and Motion Control for Multimodal Music Generation LanguageBind: Extending Video-Language Pretraining to N-modality by Language-based Semantic Alignment
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2109cd34-ce37-4fb4-ab09-158242398ce3 · outbound
EXPOTION: Facial Expression and Motion Control for Multimodal Music Generation Frechet audio distance: A reference-free metric for evaluating music enhancement algorithms,
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation b3b0825d-f5ef-4181-bdd0-f6dda41b6f6c · outbound
EXPOTION: Facial Expression and Motion Control for Multimodal Music Generation Cnn architectures for large-scale audio classification,
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 9bc0bd82-ac09-487d-ae91-6404e14bdf0a · outbound
EXPOTION: Facial Expression and Motion Control for Multimodal Music Generation Panns: Large-scale pretrained audio neural networks for audio pattern recognition,
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 4ae24de1-8228-4f7d-ad82-b95dbfabf1d2 · outbound
EXPOTION: Facial Expression and Motion Control for Multimodal Music Generation Efficient training of audio transform- ers with patchout,
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation a7efb45e-843f-4d91-96e9-5a8b69baac36 · outbound
EXPOTION: Facial Expression and Motion Control for Multimodal Music Generation Improved techniques for training gans,
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation ddce1c67-46ab-452f-9717-4bbfefbf307d · outbound
EXPOTION: Facial Expression and Motion Control for Multimodal Music Generation Large-scale contrastive language-audio pretraining with feature fusion and keyword-to-caption augmentation,
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation f793ccb1-7e6e-4408-82c8-6d38868b0d67 · outbound
EXPOTION: Facial Expression and Motion Control for Multimodal Music Generation Available: https://arxiv.org/abs/2211
Reference 2023
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 7a790e88-2b6f-4552-95e8-4e401d7c3a91 · outbound
EXPOTION: Facial Expression and Motion Control for Multimodal Music Generation Available: http://dx.doi.org/10.1016/j
Reference 2024
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9e40e093-95c3-427f-b276-9e5220945e43 · inbound
EXPOTION: Facial Expression and Motion Control for Multimodal Music Generation EXPOTION: Facial Expression and Motion Control for Multimodal Music Generation
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.