Pith. sign in

Paper Citation Record · LEDGER

EXPOTION: Facial Expression and Motion Control for Multimodal Music Generation

As of 8 August 2026, this Paper Citation Record lists 54 of 54 outbound references and 1 inbound Pith citation observation for arXiv:2507.04955.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.04955 v1

Coverage vector

measured 54 of 54 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T19:40:06.423278Z

measured 55 of 55 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T19:40:02.091028Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-06T19:40:07.168991Z

Reference resolution

54 of 54 outbound references displayed

  • verified exact2
  • verified fuzzy39
  • unresolved10
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch2

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 9e40e093-95c3-427f-b276-9e5220945e43 · outbound

This paper cites EXPOTION: Facial Expression and Motion Control for Multimodal Music Generation.

EXPOTION: Facial Expression and Motion Control for Multimodal Music Generation EXPOTION: Facial Expression and Motion Control for Multimodal Music Generation

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-08-06T19:40:07.236634Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T19:40:02.091028Z digest=sha256:3d7c060a5e6c19c7aed8f48794f31f5609690a0897401880a0129873b0adc2f4

Observation a5161bbe-ecc6-451f-9590-4f7a5a01a468 · outbound

This paper cites Visual and Motion-Based Control Early interactive sys- tems mapped facial or bodily features directly to sound.

EXPOTION: Facial Expression and Motion Control for Multimodal Music Generation Visual and Motion-Based Control Early interactive sys- tems mapped facial or bodily features directly to sound

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:40:13.582848Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T19:40:02.126591Z digest=sha256:58d093a7d7b13cbe7c82286df3afc741bfecdbe2df50b98556c9c52d7bc2c232

Observation f50d091c-f22a-40d2-8453-b09dcb9ae82a · outbound

This paper cites We froze the parame- ters of the vanilla Musicgen during training to preserve its text understanding ability.

EXPOTION: Facial Expression and Motion Control for Multimodal Music Generation We froze the parame- ters of the vanilla Musicgen during training to preserve its text understanding ability

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:40:13.439452Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T19:40:02.174306Z digest=sha256:cb43ee1f9a31cc231254e302a526932c92bb50fb058f740dcb31605b8010241f

Observation c1dde247-c637-4b46-943d-0d3213644c8a · outbound

This paper cites We recruited volunteers to record their facial expressions and upper body movements while listening to 30-second audio clips.

EXPOTION: Facial Expression and Motion Control for Multimodal Music Generation We recruited volunteers to record their facial expressions and upper body movements while listening to 30-second audio clips

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:40:13.334019Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T19:40:02.274703Z digest=sha256:3d669dbf3a32a2dc873a2c6df251adf30ea7f6891ef519a5bc9eb8865f0c634c

Observation 28781a49-92c9-4ce2-9d9a-24267fca8f1c · outbound

This paper cites an unresolved cited work.

EXPOTION: Facial Expression and Motion Control for Multimodal Music Generation Unresolved cited work

Reference 5

Resolution
unresolved
raw_fallback, observed 2026-08-06T19:40:13.173796Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T19:40:02.335175Z digest=sha256:9fea840c000fb3044bc173d16bba9112c9e5cc4c7e3db2649ff53e660c49e3d0

Observation 33c5a219-3486-401f-9e06-230fdbce7c2e · outbound

This paper cites an unresolved cited work.

EXPOTION: Facial Expression and Motion Control for Multimodal Music Generation Unresolved cited work

Reference 6

Resolution
unresolved
raw_fallback, observed 2026-08-06T19:40:13.017260Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T19:40:02.420427Z digest=sha256:e3826e8179db284479f4c2069bb557c42bea4ba5bbbc2be93ed6eefb82b37a88

Observation 5411faeb-ed22-49a3-b1a1-44ddb2381623 · outbound

This paper cites Taming Multimodal Joint Training for High-Quality Video-to-Audio Synthe- sis,.

EXPOTION: Facial Expression and Motion Control for Multimodal Music Generation Taming Multimodal Joint Training for High-Quality Video-to-Audio Synthe- sis,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:40:12.895724Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T19:40:02.506180Z digest=sha256:605a5b23129df76d081f6449f706c2730eb1c1bc97c264eef8b3e0481ce4c77a

Observation 9edde2b6-174f-4fb3-9ae6-bb3ffff98c6f · outbound

This paper cites Diff-Foley: Syn- chronized Video-to-Audio Synthesis with Latent Dif- fusion Models,.

EXPOTION: Facial Expression and Motion Control for Multimodal Music Generation Diff-Foley: Syn- chronized Video-to-Audio Synthesis with Latent Dif- fusion Models,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:40:12.771095Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T19:40:02.572574Z digest=sha256:3538615df4da2ecde0cc0054c952f38b4cd4fe3b76c512fdb61b27138928a8c2

Observation 280c36b4-c927-457d-bcaa-f567ebd5a954 · outbound

This paper cites Temporally Aligned Audio for Video with Autoregression,.

EXPOTION: Facial Expression and Motion Control for Multimodal Music Generation Temporally Aligned Audio for Video with Autoregression,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:40:12.623336Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T19:40:02.639668Z digest=sha256:a7dedeb210c6dffdd440f21e895b1cc7267bc6b6d89b8b6df7b7e3c43f49b1b5

Observation 6607a917-256c-4f2f-bd66-2a9b9f04c558 · outbound

This paper cites FoleyCrafter: Bring Silent Videos to Life with Lifelike and Synchronized Sounds.

EXPOTION: Facial Expression and Motion Control for Multimodal Music Generation FoleyCrafter: Bring Silent Videos to Life with Lifelike and Synchronized Sounds

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T19:40:02.722459Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:40:02.722459Z digest=sha256:4d04261783b2e3bbba50c9e4d6fce45950c173763667c2cc661b76e98fd8c158

Observation 7ca01989-1023-4bec-a53f-64ee01474ca4 · outbound

This paper cites Frieren: Efficient Video-to-Audio Generation Network with Rectified Flow Matching,.

EXPOTION: Facial Expression and Motion Control for Multimodal Music Generation Frieren: Efficient Video-to-Audio Generation Network with Rectified Flow Matching,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:40:12.493699Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T19:40:02.784427Z digest=sha256:99864e9dbfa3a19199260b94c1c1ad449d22df46474cf75a5a643e718f9d4a92

Observation 6d1a5a91-9a00-4e55-92d7-7bf31a0b806f · outbound

This paper cites MM-LDM: Multi-Modal Latent Diffusion Model for Sounding Video Generation,.

EXPOTION: Facial Expression and Motion Control for Multimodal Music Generation MM-LDM: Multi-Modal Latent Diffusion Model for Sounding Video Generation,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:40:12.370291Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T19:40:02.871106Z digest=sha256:209fd5eb915b8952623891ac09c5b0534104fbd1590e65f4c507234952682f22

Observation 3591a8b3-2211-42de-a461-5c3226339e1f · outbound

This paper cites Visual Echoes: A Simple Unified Transformer for Audio-Visual Generation.

EXPOTION: Facial Expression and Motion Control for Multimodal Music Generation Visual Echoes: A Simple Unified Transformer for Audio-Visual Generation

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-08-06T19:40:07.069029Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T19:40:02.934968Z digest=sha256:d37e45cd824c928645fe17fba93e4c34e128678e8b64d431c5ed432272808372

Observation 347c438d-e88c-4438-a8f0-87b421fa314a · outbound

This paper cites Video-Foley: Two-Stage Video-To-Sound Generation via Temporal Event Condition For Foley Sound,.

EXPOTION: Facial Expression and Motion Control for Multimodal Music Generation Video-Foley: Two-Stage Video-To-Sound Generation via Temporal Event Condition For Foley Sound,

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T19:40:02.986485Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:40:02.986485Z digest=sha256:a931b1d7b622eab0ae5b00bff420dcb2a55f5d1050ad3bb4e0e1b264e7e2e29a

Observation 98613fea-9f29-40f2-a7ed-af9861c212ed · outbound

This paper cites InstaGen: Enhancing Object Detection by Training on Synthetic Dataset.

EXPOTION: Facial Expression and Motion Control for Multimodal Music Generation InstaGen: Enhancing Object Detection by Training on Synthetic Dataset

Reference 15

Resolution
metadata mismatch
local_arxiv, observed 2026-08-06T19:40:06.861777Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T19:40:03.033155Z digest=sha256:43e0e06bafa85a15bda97220843b5bf908345fb71eea0c14614a4c01ae3ec5cd

Observation 1914b229-406f-4b65-95f1-44bd64a2e926 · outbound

This paper cites Orbital angular momentum of Bloch electrons: equilibrium formulation, magneto-electric phenomena, and the orbital Hall effect.

EXPOTION: Facial Expression and Motion Control for Multimodal Music Generation Orbital angular momentum of Bloch electrons: equilibrium formulation, magneto-electric phenomena, and the orbital Hall effect

Reference 16

Resolution
metadata mismatch
local_arxiv, observed 2026-08-06T19:40:06.700771Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T19:40:03.124928Z digest=sha256:78a5997ae9f2647ed64de1a5cce3b06f9322460b2384db1c4a34372dee4e216f

Observation 3268cad6-05c1-4629-991e-60b6452c29b3 · outbound

This paper cites Conditional Generation of Audio from Video via Fo- ley Analogies,.

EXPOTION: Facial Expression and Motion Control for Multimodal Music Generation Conditional Generation of Audio from Video via Fo- ley Analogies,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:40:12.236091Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T19:40:03.203009Z digest=sha256:b14b91c5d53c1b309328c4895090a55ac4d343d1723a3ee7a931ac1956679e7c

Observation 570ab3ab-09c6-4f44-a80b-eb1002082a44 · outbound

This paper cites MuVi: Video-to-Music Generation with Semantic Alignment and Rhythmic Synchronization.

EXPOTION: Facial Expression and Motion Control for Multimodal Music Generation MuVi: Video-to-Music Generation with Semantic Alignment and Rhythmic Synchronization

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T19:40:03.310979Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:40:03.310979Z digest=sha256:57e2cffce7cf894e9fdbe29b8366f5567e85bfe8163f9b0b1c0e5cda2d397b40

Observation 5c4325b2-3013-4bf5-94e1-01675403956d · outbound

This paper cites VMAs: Video-to-Music Generation via Semantic Alignment in Web Music Videos,.

EXPOTION: Facial Expression and Motion Control for Multimodal Music Generation VMAs: Video-to-Music Generation via Semantic Alignment in Web Music Videos,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:40:12.123750Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T19:40:03.401713Z digest=sha256:2df5b5a051e699ea2c7e8d36c05dd29c2431cd7fe3f8bc859a56d11d0cff6d1b

Observation ace31055-2ce3-4c57-9ada-33435a441f6d · outbound

This paper cites Video2Music: Suitable Music Generation from Videos using an Affec- tive Multimodal Transformer model,.

EXPOTION: Facial Expression and Motion Control for Multimodal Music Generation Video2Music: Suitable Music Generation from Videos using an Affec- tive Multimodal Transformer model,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:40:11.993773Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T19:40:03.480431Z digest=sha256:61aab485fc84053c11c18bf669ed2eaa316165c1d6da0fcd024bb4c3c375fa53

Observation 272d2f30-f92d-4de2-88de-12a9bd7183b7 · outbound

This paper cites V2Meow: Meowing to the Visual Beat via Video-to-Music Generation,.

EXPOTION: Facial Expression and Motion Control for Multimodal Music Generation V2Meow: Meowing to the Visual Beat via Video-to-Music Generation,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:40:11.847780Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T19:40:03.578377Z digest=sha256:8a37f44a33b734f01a4e16fa41bbcfa73ff3d8c2b54f6666c94c46d5460f2a7a

Observation df50cb85-48d6-4390-8928-3e959bdb753e · outbound

This paper cites DanceCom- poser: Dance-to-Music Generation Using a Progressive Conditional Music Generator,.

EXPOTION: Facial Expression and Motion Control for Multimodal Music Generation DanceCom- poser: Dance-to-Music Generation Using a Progressive Conditional Music Generator,

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:40:11.704582Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T19:40:03.686152Z digest=sha256:95bab25a4889b9bb91c4aada565fff9fa017892c03a5415af7e2b8055b122865

Observation 7e4af507-c12d-48d4-ad6b-2f3f28224bbe · outbound

This paper cites Discrete Contrastive Diffusion for Cross- Modal Music and Image Generation,.

EXPOTION: Facial Expression and Motion Control for Multimodal Music Generation Discrete Contrastive Diffusion for Cross- Modal Music and Image Generation,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:40:11.528547Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T19:40:03.792531Z digest=sha256:334b650bd285ccf35a8eccc55d3e54fac724ffdf4e649aab37213b70d86e35b4

Observation 7c3a79c2-7f4f-4cd3-a3e4-95985327a9f3 · outbound

This paper cites Long- Term Rhythmic Video Soundtracker,.

EXPOTION: Facial Expression and Motion Control for Multimodal Music Generation Long- Term Rhythmic Video Soundtracker,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:40:11.368600Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T19:40:03.912137Z digest=sha256:6419e88663f3e08571badd383e92c0028237648c9ced457447c0290410fcd0b5

Observation 20b4b726-aaae-4943-aaff-ce82af8fc7c4 · outbound

This paper cites Simple and control- lable music generation,.

EXPOTION: Facial Expression and Motion Control for Multimodal Music Generation Simple and control- lable music generation,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:40:11.195413Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T19:40:03.992590Z digest=sha256:646397c24d50ce2a81415222a70a2d68d56b6cde924b883cdf22fa6ca7ba3f4b

Observation 8f0bba4b-2550-4491-86ea-2112d40c1c49 · outbound

This paper cites Content-based Controls For Music Large Language Modeling.

EXPOTION: Facial Expression and Motion Control for Multimodal Music Generation Content-based Controls For Music Large Language Modeling

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T19:40:04.068660Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:40:04.068660Z digest=sha256:97f1e8ca01aace21973066bfa31dbdd9d3d8d0a866af11246bbe0f655b48acdf

Observation 867e4e4d-3533-4f57-9782-533fc52b2236 · outbound

This paper cites BiMediX: Bilingual Medical Mixture of Experts LLM.

EXPOTION: Facial Expression and Motion Control for Multimodal Music Generation BiMediX: Bilingual Medical Mixture of Experts LLM

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T19:40:04.143427Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:40:04.143427Z digest=sha256:8fecad5770da1b579fdf9d3cffb97d5544cf95d1a0c83c468c42a45c56263eda

Observation 95569952-453b-4786-a063-b217798b55b4 · outbound

This paper cites Audioclip: Extending clip to image, text and audio,.

EXPOTION: Facial Expression and Motion Control for Multimodal Music Generation Audioclip: Extending clip to image, text and audio,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:40:11.013911Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T19:40:04.196472Z digest=sha256:e53d03c67688281296c87ca31cbc42edd63ac57a3c9959f554335c8465b29b28

Observation 660763c1-7f85-49a8-abe1-d7d88733bbf8 · outbound

This paper cites Exploring the limits of transfer learning with a unified text-to-text transformer,.

EXPOTION: Facial Expression and Motion Control for Multimodal Music Generation Exploring the limits of transfer learning with a unified text-to-text transformer,

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T19:40:04.281404Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:40:04.281404Z digest=sha256:0f8f05ba62c2a2bae8f08d73cef6903ac6c8328218fca767b2b110e0e36bb309

Observation 1a77d445-28de-4259-85b5-7470c0e565af · outbound

This paper cites Vidmuse: A simple video-to- music generation framework with long-short-term mod- eling,.

EXPOTION: Facial Expression and Motion Control for Multimodal Music Generation Vidmuse: A simple video-to- music generation framework with long-short-term mod- eling,

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:40:10.832849Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T19:40:04.379020Z digest=sha256:cdd7a769c519a62e095e688e313c3703ebe75d4f2038c087978c6fc4b3afd12a

Observation b3655a34-319f-49d9-8e80-239d6fc3e059 · outbound

This paper cites Sonify Your Face: Facial Expressions for Sound Generation,.

EXPOTION: Facial Expression and Motion Control for Multimodal Music Generation Sonify Your Face: Facial Expressions for Sound Generation,

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:40:10.626318Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T19:40:04.446437Z digest=sha256:0c1151ce04ea4e5c852f36bcada95d698ce57f54bf9464c1c3a6e9eb7b883410

Observation 213f40a7-211f-4da6-aeae-8ee1e726f045 · outbound

This paper cites Movement to emotions to music: using whole body emotional expression as an interaction for electronic music genera- tion,.

EXPOTION: Facial Expression and Motion Control for Multimodal Music Generation Movement to emotions to music: using whole body emotional expression as an interaction for electronic music genera- tion,

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:40:10.373609Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T19:40:04.532491Z digest=sha256:0f6049b8a85641fbceb5e20c3ef1049ba6ca3f937be10fb427b1a347c55c7799

Observation 5a30467f-308d-41b2-b8e5-4398e4ffedfb · outbound

This paper cites D2MNet for music generation jointly driven by facial expressions and dance movements,.

EXPOTION: Facial Expression and Motion Control for Multimodal Music Generation D2MNet for music generation jointly driven by facial expressions and dance movements,

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:40:10.145199Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T19:40:04.603651Z digest=sha256:6bc28b6787183f6b1bbf5d4246e2e901097c6264494568c0fbdba81605a27f02

Observation 8beac812-bcff-45fc-9d2a-dfb7b735dc7c · outbound

This paper cites Deep- Tunes: Music Generation based on Facial Emotions using Deep Learning,.

EXPOTION: Facial Expression and Motion Control for Multimodal Music Generation Deep- Tunes: Music Generation based on Facial Emotions using Deep Learning,

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:40:09.946722Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T19:40:04.683328Z digest=sha256:a23f08c0ac90cb9f78da8a7c3507132c212f35112ea29017b541cc4daf2266d6

Observation 95e5e6ef-e7cd-4911-b0e4-dd7226c2d195 · outbound

This paper cites A Contin- uous Emotional Music Generation System Based on Facial Expressions,.

EXPOTION: Facial Expression and Motion Control for Multimodal Music Generation A Contin- uous Emotional Music Generation System Based on Facial Expressions,

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:40:09.797652Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T19:40:04.816135Z digest=sha256:6fc783efaf23ac13baa6debf0a929976668f5a58ae4e4453579a3d50601ee88d

Observation 290f3bdf-e9f3-492b-8641-be59b2b62b56 · outbound

This paper cites Musiclm: Generating music from text,.

EXPOTION: Facial Expression and Motion Control for Multimodal Music Generation Musiclm: Generating music from text,

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:40:09.655800Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T19:40:04.908406Z digest=sha256:663bd61bb7abe6374f045240dd1cf47cef18bbb020dbd4e781d692a51a2da932

Observation 437cbddc-94b7-4e31-b553-70831a24277a · outbound

This paper cites Audio set: An on- tology and human-labeled dataset for audio events,.

EXPOTION: Facial Expression and Motion Control for Multimodal Music Generation Audio set: An on- tology and human-labeled dataset for audio events,

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:40:09.502002Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T19:40:05.007593Z digest=sha256:4a25976c3500359dec36c8849a5d26e90a8ab71ee41ffb58d74aee8ac97084e4

Observation 8d4a48bf-9cba-4a3e-a001-b3488749bedb · outbound

This paper cites Vg- gsound: A large-scale audio-visual dataset,.

EXPOTION: Facial Expression and Motion Control for Multimodal Music Generation Vg- gsound: A large-scale audio-visual dataset,

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:40:09.346892Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T19:40:05.092253Z digest=sha256:6b7d65282653290d3599acaadc1baaf9786f8662d9c023f45ab84056e67792ec

Observation 22542eab-239f-45f9-a1d1-3db28816f211 · outbound

This paper cites Arrange, inpaint, and refine: Steerable long-term music audio generation and editing via content-based controls,.

EXPOTION: Facial Expression and Motion Control for Multimodal Music Generation Arrange, inpaint, and refine: Steerable long-term music audio generation and editing via content-based controls,

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:40:09.182866Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T19:40:05.158914Z digest=sha256:5fdc9c3f9dd0b97df4d98411e1c50fb51b1519dde890ae34a5a5badab58df824

Observation c29f1b56-bd19-4960-bb1b-e28fac85f60b · outbound

This paper cites Marlin: Masked autoencoder for facial video representation learning,.

EXPOTION: Facial Expression and Motion Control for Multimodal Music Generation Marlin: Masked autoencoder for facial video representation learning,

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:40:09.043876Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T19:40:05.230502Z digest=sha256:0f12284022a091b0f929570bc1aeaa325f4eed36db733e0619653c99d0c6d64f

Observation 83659852-de37-4dcc-a2fc-0941b0e9eaea · outbound

This paper cites Synch- former: Efficient synchronization from sparse cues,.

EXPOTION: Facial Expression and Motion Control for Multimodal Music Generation Synch- former: Efficient synchronization from sparse cues,

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:40:08.910894Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T19:40:05.317236Z digest=sha256:6b79e665577430ddfa8c72fd27c3508e98acbc396fcdf11bdf7e3616bb93b112

Observation 8ee1745e-fb97-48c2-8d7c-8b64bbcdc768 · outbound

This paper cites Raft: Recurrent all-pairs field transforms for optical flow,.

EXPOTION: Facial Expression and Motion Control for Multimodal Music Generation Raft: Recurrent all-pairs field transforms for optical flow,

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:40:08.790842Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T19:40:05.387260Z digest=sha256:9fc544bfec5e53cba4547cb1cf4e7f91af037295ac2c1bda304b224a18113688

Observation 30df38f1-ba3e-4b42-9166-3b4bc2c8f2e6 · outbound

This paper cites Salmonn: Towards generic hear- ing abilities for large language models,.

EXPOTION: Facial Expression and Motion Control for Multimodal Music Generation Salmonn: Towards generic hear- ing abilities for large language models,

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:40:08.638745Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T19:40:05.475020Z digest=sha256:befcb8cba9e272c1a72a7cf57e47bd88f30b6687d3a2df4316dbfd7b639aeb7c

Observation a906a825-ceb7-4767-a58c-469be428b1c0 · outbound

This paper cites High Fidelity Neural Audio Compression.

EXPOTION: Facial Expression and Motion Control for Multimodal Music Generation High Fidelity Neural Audio Compression

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-06T19:40:05.562715Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:40:05.562715Z digest=sha256:5abf933e50e5a1a4f013cb760fa5bc9f1b4cd772cc1d777530adff7549efbda9

Observation 3570c2dd-fc12-4839-94ba-a278eac465da · outbound

This paper cites Video2music: Suitable music generation from videos using an affective multimodal transformer model,.

EXPOTION: Facial Expression and Motion Control for Multimodal Music Generation Video2music: Suitable music generation from videos using an affective multimodal transformer model,

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:40:08.478892Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T19:40:05.646574Z digest=sha256:c2f70015b65ea7b1dd9c9471075e1c31fede47e0194791ca473a33ed70ada2d1

Observation 4d498afb-aabf-454d-b1f7-7cb6a779273d · outbound

This paper cites LanguageBind: Extending Video-Language Pretraining to N-modality by Language-based Semantic Alignment.

EXPOTION: Facial Expression and Motion Control for Multimodal Music Generation LanguageBind: Extending Video-Language Pretraining to N-modality by Language-based Semantic Alignment

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-06T19:40:06.423278Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:40:06.423278Z digest=sha256:4099d7afdc6e17bf5ca8bbb54cb37e1b73a807602acb34f7e960ec4593194f1a

Observation 2109cd34-ce37-4fb4-ab09-158242398ce3 · outbound

This paper cites Frechet audio distance: A reference-free metric for evaluating music enhancement algorithms,.

EXPOTION: Facial Expression and Motion Control for Multimodal Music Generation Frechet audio distance: A reference-free metric for evaluating music enhancement algorithms,

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:40:08.324059Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T19:40:05.802192Z digest=sha256:80839b66d27e086ee4cdb654382fcbe4e0463b7375e485644191238b0f5f0bb8

Observation b3b0825d-f5ef-4181-bdd0-f6dda41b6f6c · outbound

This paper cites Cnn architectures for large-scale audio classification,.

EXPOTION: Facial Expression and Motion Control for Multimodal Music Generation Cnn architectures for large-scale audio classification,

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:40:08.184873Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T19:40:05.915584Z digest=sha256:43c0399eecc22a3f111ff86ecd28c8c86a67e80d1ef58013c5113251824a0c68

Observation 9bc0bd82-ac09-487d-ae91-6404e14bdf0a · outbound

This paper cites Panns: Large-scale pretrained audio neural networks for audio pattern recognition,.

EXPOTION: Facial Expression and Motion Control for Multimodal Music Generation Panns: Large-scale pretrained audio neural networks for audio pattern recognition,

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:40:08.043067Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T19:40:06.008469Z digest=sha256:c68dca8fa6fa25e8afe7f8a6f9c1b7086eb3f110fab00fadcbb0893deb9b9877

Observation 4ae24de1-8228-4f7d-ad82-b95dbfabf1d2 · outbound

This paper cites Efficient training of audio transform- ers with patchout,.

EXPOTION: Facial Expression and Motion Control for Multimodal Music Generation Efficient training of audio transform- ers with patchout,

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:40:07.890135Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T19:40:06.095672Z digest=sha256:c8eaa9a4c7a24aa0af3b8814ac3f93613b4b125ccace6b4a29eea972730527c7

Observation a7efb45e-843f-4d91-96e9-5a8b69baac36 · outbound

This paper cites Improved techniques for training gans,.

EXPOTION: Facial Expression and Motion Control for Multimodal Music Generation Improved techniques for training gans,

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:40:07.747301Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T19:40:06.183184Z digest=sha256:a902a73c2026a9e2899abb25a63a8677cd29fb0397cc6c503950145285f8b3c9

Observation ddce1c67-46ab-452f-9717-4bbfefbf307d · outbound

This paper cites Large-scale contrastive language-audio pretraining with feature fusion and keyword-to-caption augmentation,.

EXPOTION: Facial Expression and Motion Control for Multimodal Music Generation Large-scale contrastive language-audio pretraining with feature fusion and keyword-to-caption augmentation,

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:40:07.534298Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T19:40:06.270020Z digest=sha256:113c8034f9d83b97e1b7121942328058af30637f49f514862d321516c9b38b73

Observation f793ccb1-7e6e-4408-82c8-6d38868b0d67 · outbound

This paper cites Available: https://arxiv.org/abs/2211.

EXPOTION: Facial Expression and Motion Control for Multimodal Music Generation Available: https://arxiv.org/abs/2211

Reference 2023

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:40:07.382378Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T19:40:06.341162Z digest=sha256:e265e29799ab77c3c95864a1d4727e7f26afcdd9b40b43fb0897ec63bcb361ea

Observation 7a790e88-2b6f-4552-95e8-4e401d7c3a91 · outbound

This paper cites Available: http://dx.doi.org/10.1016/j.

EXPOTION: Facial Expression and Motion Control for Multimodal Music Generation Available: http://dx.doi.org/10.1016/j

Reference 2024

Resolution
malformed identifier
no resolver link, observed 2026-08-06T19:40:05.719311Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:40:05.719311Z digest=sha256:6d0a87dd838740f94b8c4fe20a5809afe273da61e83384ed15fa23fcac7b421b

Pith citing papers

Observation 9e40e093-95c3-427f-b276-9e5220945e43 · inbound

EXPOTION: Facial Expression and Motion Control for Multimodal Music Generation cites this paper.

EXPOTION: Facial Expression and Motion Control for Multimodal Music Generation EXPOTION: Facial Expression and Motion Control for Multimodal Music Generation

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-08-06T19:40:07.236634Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T19:40:02.091028Z digest=sha256:3d7c060a5e6c19c7aed8f48794f31f5609690a0897401880a0129873b0adc2f4