Pith. sign in

Paper Citation Record · LEDGER

EXPOTION: Facial Expression and Motion Control for Multimodal Music Generation

As of 21 August 2026, this Paper Citation Record lists 54 of 54 outbound references and 1 inbound Pith citation observation for arXiv:2507.04955.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.04955 v1

Coverage vector

measured 54 of 54 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T19:40:06.423278Z

measured 55 of 55 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T19:40:02.091028Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-06T19:40:07.168991Z

Reference resolution

54 of 54 outbound references displayed

  • verified exact2
  • verified fuzzy39
  • unresolved10
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch2

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 9e40e093-95c3-427f-b276-9e5220945e43 · outbound

This paper cites EXPOTION: Facial Expression and Motion Control for Multimodal Music Generation.

EXPOTION: Facial Expression and Motion Control for Multimodal Music Generation EXPOTION: Facial Expression and Motion Control for Multimodal Music Generation

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-08-06T19:40:07.236634Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T19:40:02.091028Z digest=sha256:5fe1c0dd48636462e9b5ad5754dec57c13d102b054e81588de26f0363f9dcbd5

Observation a5161bbe-ecc6-451f-9590-4f7a5a01a468 · outbound

This paper cites Visual and Motion-Based Control Early interactive sys- tems mapped facial or bodily features directly to sound.

EXPOTION: Facial Expression and Motion Control for Multimodal Music Generation Visual and Motion-Based Control Early interactive sys- tems mapped facial or bodily features directly to sound

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:40:13.582848Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T19:40:02.126591Z digest=sha256:61d54f1304ddab9b446e2a14f316d8e7bcd348ff02cb640a0d79d83e4e2e6712

Observation f50d091c-f22a-40d2-8453-b09dcb9ae82a · outbound

This paper cites We froze the parame- ters of the vanilla Musicgen during training to preserve its text understanding ability.

EXPOTION: Facial Expression and Motion Control for Multimodal Music Generation We froze the parame- ters of the vanilla Musicgen during training to preserve its text understanding ability

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:40:13.439452Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T19:40:02.174306Z digest=sha256:37324739599a0aefdaea8d292fd9d4e8b6299c44135b36d27df72a4bf5a8ad13

Observation c1dde247-c637-4b46-943d-0d3213644c8a · outbound

This paper cites We recruited volunteers to record their facial expressions and upper body movements while listening to 30-second audio clips.

EXPOTION: Facial Expression and Motion Control for Multimodal Music Generation We recruited volunteers to record their facial expressions and upper body movements while listening to 30-second audio clips

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:40:13.334019Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T19:40:02.274703Z digest=sha256:2387e42c1865251a470dd3c0e677cce1d9b2cd6467b285619d24a10bede038c9

Observation 28781a49-92c9-4ce2-9d9a-24267fca8f1c · outbound

This paper cites an unresolved cited work.

EXPOTION: Facial Expression and Motion Control for Multimodal Music Generation Unresolved cited work

Reference 5

Resolution
unresolved
raw_fallback, observed 2026-08-06T19:40:13.173796Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T19:40:02.335175Z digest=sha256:8d58f9abb363b1a8b7120dac149678185272fa8de7faa69be981d2da3cedc8c8

Observation 33c5a219-3486-401f-9e06-230fdbce7c2e · outbound

This paper cites an unresolved cited work.

EXPOTION: Facial Expression and Motion Control for Multimodal Music Generation Unresolved cited work

Reference 6

Resolution
unresolved
raw_fallback, observed 2026-08-06T19:40:13.017260Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T19:40:02.420427Z digest=sha256:20218b2fe1dc518b2c631009244fa87a8ba7379c75988dd94b9594df6726b3a1

Observation 5411faeb-ed22-49a3-b1a1-44ddb2381623 · outbound

This paper cites Taming Multimodal Joint Training for High-Quality Video-to-Audio Synthe- sis,.

EXPOTION: Facial Expression and Motion Control for Multimodal Music Generation Taming Multimodal Joint Training for High-Quality Video-to-Audio Synthe- sis,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:40:12.895724Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T19:40:02.506180Z digest=sha256:5167ede00cb806828cc6e2e1fe103dcbe04a213baf93e8ed67429f77d7338acd

Observation 9edde2b6-174f-4fb3-9ae6-bb3ffff98c6f · outbound

This paper cites Diff-Foley: Syn- chronized Video-to-Audio Synthesis with Latent Dif- fusion Models,.

EXPOTION: Facial Expression and Motion Control for Multimodal Music Generation Diff-Foley: Syn- chronized Video-to-Audio Synthesis with Latent Dif- fusion Models,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:40:12.771095Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T19:40:02.572574Z digest=sha256:11373ac433dd8442ded1c59fbb72b77d5a8394066313bc5d863349c9196dc325

Observation 280c36b4-c927-457d-bcaa-f567ebd5a954 · outbound

This paper cites Temporally Aligned Audio for Video with Autoregression,.

EXPOTION: Facial Expression and Motion Control for Multimodal Music Generation Temporally Aligned Audio for Video with Autoregression,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:40:12.623336Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T19:40:02.639668Z digest=sha256:7c1caf39e86ef7cfcdead8e9a988f44492777f224e44ca366e4579722aff931e

Observation 6607a917-256c-4f2f-bd66-2a9b9f04c558 · outbound

This paper cites FoleyCrafter: Bring Silent Videos to Life with Lifelike and Synchronized Sounds.

EXPOTION: Facial Expression and Motion Control for Multimodal Music Generation FoleyCrafter: Bring Silent Videos to Life with Lifelike and Synchronized Sounds

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T19:40:02.722459Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:40:02.722459Z digest=sha256:7c3dc51acf4e433d1f62b37a7cff972e326ee6c1f4ca93457ab16028e8dcf5d3

Observation 7ca01989-1023-4bec-a53f-64ee01474ca4 · outbound

This paper cites Frieren: Efficient Video-to-Audio Generation Network with Rectified Flow Matching,.

EXPOTION: Facial Expression and Motion Control for Multimodal Music Generation Frieren: Efficient Video-to-Audio Generation Network with Rectified Flow Matching,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:40:12.493699Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T19:40:02.784427Z digest=sha256:baf3e089c066cfafb003a3b4aa7235208c70c23b73895353fb91b0664192e051

Observation 6d1a5a91-9a00-4e55-92d7-7bf31a0b806f · outbound

This paper cites MM-LDM: Multi-Modal Latent Diffusion Model for Sounding Video Generation,.

EXPOTION: Facial Expression and Motion Control for Multimodal Music Generation MM-LDM: Multi-Modal Latent Diffusion Model for Sounding Video Generation,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:40:12.370291Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T19:40:02.871106Z digest=sha256:30a2112afbf2b3648000ec2df34f4688aa2e4af529b4d1db63ed70f8cdbb66c7

Observation 3591a8b3-2211-42de-a461-5c3226339e1f · outbound

This paper cites Visual Echoes: A Simple Unified Transformer for Audio-Visual Generation.

EXPOTION: Facial Expression and Motion Control for Multimodal Music Generation Visual Echoes: A Simple Unified Transformer for Audio-Visual Generation

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-08-06T19:40:07.069029Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T19:40:02.934968Z digest=sha256:70124ce6ecdc11dd99f494d88f017dd1e8621303aa922af33d1dfa4cf42009d0

Observation 347c438d-e88c-4438-a8f0-87b421fa314a · outbound

This paper cites Video-Foley: Two-Stage Video-To-Sound Generation via Temporal Event Condition For Foley Sound,.

EXPOTION: Facial Expression and Motion Control for Multimodal Music Generation Video-Foley: Two-Stage Video-To-Sound Generation via Temporal Event Condition For Foley Sound,

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T19:40:02.986485Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:40:02.986485Z digest=sha256:ffa8b622cd38f476c7d8f06fa0ebd6fc611fe8b54bb9109f99b021721d031d44

Observation 98613fea-9f29-40f2-a7ed-af9861c212ed · outbound

This paper cites InstaGen: Enhancing Object Detection by Training on Synthetic Dataset.

EXPOTION: Facial Expression and Motion Control for Multimodal Music Generation InstaGen: Enhancing Object Detection by Training on Synthetic Dataset

Reference 15

Resolution
metadata mismatch
local_arxiv, observed 2026-08-06T19:40:06.861777Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T19:40:03.033155Z digest=sha256:8f9c9630deeb51b49da171185e182d24542bc49669fc7e789a4595c670b643f5

Observation 1914b229-406f-4b65-95f1-44bd64a2e926 · outbound

This paper cites Orbital angular momentum of Bloch electrons: equilibrium formulation, magneto-electric phenomena, and the orbital Hall effect.

EXPOTION: Facial Expression and Motion Control for Multimodal Music Generation Orbital angular momentum of Bloch electrons: equilibrium formulation, magneto-electric phenomena, and the orbital Hall effect

Reference 16

Resolution
metadata mismatch
local_arxiv, observed 2026-08-06T19:40:06.700771Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T19:40:03.124928Z digest=sha256:f4fdb49fc81f24e4a77f970d3d7b1650ebfecc40ebad0d142d68d275e06bd2b4

Observation 3268cad6-05c1-4629-991e-60b6452c29b3 · outbound

This paper cites Conditional Generation of Audio from Video via Fo- ley Analogies,.

EXPOTION: Facial Expression and Motion Control for Multimodal Music Generation Conditional Generation of Audio from Video via Fo- ley Analogies,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:40:12.236091Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T19:40:03.203009Z digest=sha256:f08b7c4a4727e971084e5f65db4ee9ccd6ff27f5dd1ea6847b94f91ea010f3c4

Observation 570ab3ab-09c6-4f44-a80b-eb1002082a44 · outbound

This paper cites MuVi: Video-to-Music Generation with Semantic Alignment and Rhythmic Synchronization.

EXPOTION: Facial Expression and Motion Control for Multimodal Music Generation MuVi: Video-to-Music Generation with Semantic Alignment and Rhythmic Synchronization

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T19:40:03.310979Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:40:03.310979Z digest=sha256:789131dd798a6b3f33314af9676ce9315247c09f232de97f5934b2ff1cc4f889

Observation 5c4325b2-3013-4bf5-94e1-01675403956d · outbound

This paper cites VMAs: Video-to-Music Generation via Semantic Alignment in Web Music Videos,.

EXPOTION: Facial Expression and Motion Control for Multimodal Music Generation VMAs: Video-to-Music Generation via Semantic Alignment in Web Music Videos,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:40:12.123750Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T19:40:03.401713Z digest=sha256:2e2d8506c629c15363213e406cd098d3c7384674fd89284e6cefd9ae55c97cd3

Observation ace31055-2ce3-4c57-9ada-33435a441f6d · outbound

This paper cites Video2Music: Suitable Music Generation from Videos using an Affec- tive Multimodal Transformer model,.

EXPOTION: Facial Expression and Motion Control for Multimodal Music Generation Video2Music: Suitable Music Generation from Videos using an Affec- tive Multimodal Transformer model,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:40:11.993773Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T19:40:03.480431Z digest=sha256:8b472f339d426b60c7bd87b6b484eec96ddb8d0edec74e36344ebd76e7c5e747

Observation 272d2f30-f92d-4de2-88de-12a9bd7183b7 · outbound

This paper cites V2Meow: Meowing to the Visual Beat via Video-to-Music Generation,.

EXPOTION: Facial Expression and Motion Control for Multimodal Music Generation V2Meow: Meowing to the Visual Beat via Video-to-Music Generation,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:40:11.847780Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T19:40:03.578377Z digest=sha256:96d5a8ca8f697fffe92405e94d1dd5f4d29eba074b601bb63e3431d968dd28ce

Observation df50cb85-48d6-4390-8928-3e959bdb753e · outbound

This paper cites DanceCom- poser: Dance-to-Music Generation Using a Progressive Conditional Music Generator,.

EXPOTION: Facial Expression and Motion Control for Multimodal Music Generation DanceCom- poser: Dance-to-Music Generation Using a Progressive Conditional Music Generator,

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:40:11.704582Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T19:40:03.686152Z digest=sha256:938a3c7b41368517aceb6079964c1e30c5e424b9ec4d9a236b4e320b4f8d3a44

Observation 7e4af507-c12d-48d4-ad6b-2f3f28224bbe · outbound

This paper cites Discrete Contrastive Diffusion for Cross- Modal Music and Image Generation,.

EXPOTION: Facial Expression and Motion Control for Multimodal Music Generation Discrete Contrastive Diffusion for Cross- Modal Music and Image Generation,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:40:11.528547Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T19:40:03.792531Z digest=sha256:bed07d233c906f5c53e0db297ddb1ba013119900108d078a6097fbd0af56e5f3

Observation 7c3a79c2-7f4f-4cd3-a3e4-95985327a9f3 · outbound

This paper cites Long- Term Rhythmic Video Soundtracker,.

EXPOTION: Facial Expression and Motion Control for Multimodal Music Generation Long- Term Rhythmic Video Soundtracker,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:40:11.368600Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T19:40:03.912137Z digest=sha256:19a80b4f9bb8b656d52b3000faf71591ed9f958f86702a495481e65d31c94e5b

Observation 20b4b726-aaae-4943-aaff-ce82af8fc7c4 · outbound

This paper cites Simple and control- lable music generation,.

EXPOTION: Facial Expression and Motion Control for Multimodal Music Generation Simple and control- lable music generation,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:40:11.195413Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T19:40:03.992590Z digest=sha256:991dea5f3b88dae8a469e8a8172abea36134f6c035f5d56fbf35db0abd621e9e

Observation 8f0bba4b-2550-4491-86ea-2112d40c1c49 · outbound

This paper cites Content-based Controls For Music Large Language Modeling.

EXPOTION: Facial Expression and Motion Control for Multimodal Music Generation Content-based Controls For Music Large Language Modeling

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T19:40:04.068660Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:40:04.068660Z digest=sha256:6d9ee4f53cd2f8f84831f026aeba05b37572fc03851d32d8cd1f6d8acc969cac

Observation 867e4e4d-3533-4f57-9782-533fc52b2236 · outbound

This paper cites BiMediX: Bilingual Medical Mixture of Experts LLM.

EXPOTION: Facial Expression and Motion Control for Multimodal Music Generation BiMediX: Bilingual Medical Mixture of Experts LLM

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T19:40:04.143427Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:40:04.143427Z digest=sha256:43ddfe065421030111993a4cf8f617f25d8cbf4e8de08c6c7e72d486579b926c

Observation 95569952-453b-4786-a063-b217798b55b4 · outbound

This paper cites Audioclip: Extending clip to image, text and audio,.

EXPOTION: Facial Expression and Motion Control for Multimodal Music Generation Audioclip: Extending clip to image, text and audio,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:40:11.013911Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T19:40:04.196472Z digest=sha256:e17e6e5a68e8bcea03063ff19e403d319fec07f10c7ee0e168e33991df695c6f

Observation 660763c1-7f85-49a8-abe1-d7d88733bbf8 · outbound

This paper cites Exploring the limits of transfer learning with a unified text-to-text transformer,.

EXPOTION: Facial Expression and Motion Control for Multimodal Music Generation Exploring the limits of transfer learning with a unified text-to-text transformer,

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T19:40:04.281404Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:40:04.281404Z digest=sha256:0117249a41a2630ae3f481d1446b62cce591d76c7a9bed5e905434881500b4bd

Observation 1a77d445-28de-4259-85b5-7470c0e565af · outbound

This paper cites Vidmuse: A simple video-to- music generation framework with long-short-term mod- eling,.

EXPOTION: Facial Expression and Motion Control for Multimodal Music Generation Vidmuse: A simple video-to- music generation framework with long-short-term mod- eling,

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:40:10.832849Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T19:40:04.379020Z digest=sha256:3da6deecb7c08820245a4ace66aa56fd09131d6b1c9e4c0d72b7fbfa1a5d282c

Observation b3655a34-319f-49d9-8e80-239d6fc3e059 · outbound

This paper cites Sonify Your Face: Facial Expressions for Sound Generation,.

EXPOTION: Facial Expression and Motion Control for Multimodal Music Generation Sonify Your Face: Facial Expressions for Sound Generation,

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:40:10.626318Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T19:40:04.446437Z digest=sha256:e88914ebc09e2ce7ce49901e5d43427d3c3b75d1c05f29caa1b4e98d8a481eab

Observation 213f40a7-211f-4da6-aeae-8ee1e726f045 · outbound

This paper cites Movement to emotions to music: using whole body emotional expression as an interaction for electronic music genera- tion,.

EXPOTION: Facial Expression and Motion Control for Multimodal Music Generation Movement to emotions to music: using whole body emotional expression as an interaction for electronic music genera- tion,

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:40:10.373609Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T19:40:04.532491Z digest=sha256:5a548829b8789b733d463c223b8d19561dee72d5221a236f60a89bd1fff63929

Observation 5a30467f-308d-41b2-b8e5-4398e4ffedfb · outbound

This paper cites D2MNet for music generation jointly driven by facial expressions and dance movements,.

EXPOTION: Facial Expression and Motion Control for Multimodal Music Generation D2MNet for music generation jointly driven by facial expressions and dance movements,

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:40:10.145199Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T19:40:04.603651Z digest=sha256:85c110bbae1bff3eb3f81547c78a49cc17f5c412d2a6dec4dc3ba7886f31535d

Observation 8beac812-bcff-45fc-9d2a-dfb7b735dc7c · outbound

This paper cites Deep- Tunes: Music Generation based on Facial Emotions using Deep Learning,.

EXPOTION: Facial Expression and Motion Control for Multimodal Music Generation Deep- Tunes: Music Generation based on Facial Emotions using Deep Learning,

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:40:09.946722Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T19:40:04.683328Z digest=sha256:9e9b7814e4cef0dba46422831969afdbacdc538068c9c2cede2a1e04e4c62bda

Observation 95e5e6ef-e7cd-4911-b0e4-dd7226c2d195 · outbound

This paper cites A Contin- uous Emotional Music Generation System Based on Facial Expressions,.

EXPOTION: Facial Expression and Motion Control for Multimodal Music Generation A Contin- uous Emotional Music Generation System Based on Facial Expressions,

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:40:09.797652Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T19:40:04.816135Z digest=sha256:2543bde13d89b36e5d971ffa52640851471fc5efe4733e62c91318951ccb5cdd

Observation 290f3bdf-e9f3-492b-8641-be59b2b62b56 · outbound

This paper cites Musiclm: Generating music from text,.

EXPOTION: Facial Expression and Motion Control for Multimodal Music Generation Musiclm: Generating music from text,

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:40:09.655800Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T19:40:04.908406Z digest=sha256:f88043b14627faec4cd52fcafafcd4ccbfcc3fc0a06687f186f3276c46711b19

Observation 437cbddc-94b7-4e31-b553-70831a24277a · outbound

This paper cites Audio set: An on- tology and human-labeled dataset for audio events,.

EXPOTION: Facial Expression and Motion Control for Multimodal Music Generation Audio set: An on- tology and human-labeled dataset for audio events,

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:40:09.502002Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T19:40:05.007593Z digest=sha256:0cf309d1954214a0a11f951d98270bbcfaf32a8444f74a09d7dd65d644c41f2e

Observation 8d4a48bf-9cba-4a3e-a001-b3488749bedb · outbound

This paper cites Vg- gsound: A large-scale audio-visual dataset,.

EXPOTION: Facial Expression and Motion Control for Multimodal Music Generation Vg- gsound: A large-scale audio-visual dataset,

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:40:09.346892Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T19:40:05.092253Z digest=sha256:2ca21fad611b212f7008d780c0dcd23a040afd7e5c3568ba046ed170005944f7

Observation 22542eab-239f-45f9-a1d1-3db28816f211 · outbound

This paper cites Arrange, inpaint, and refine: Steerable long-term music audio generation and editing via content-based controls,.

EXPOTION: Facial Expression and Motion Control for Multimodal Music Generation Arrange, inpaint, and refine: Steerable long-term music audio generation and editing via content-based controls,

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:40:09.182866Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T19:40:05.158914Z digest=sha256:339d5e42b3e37c122dce0631b31054a962a8e2e9928f5876274d769164dfb1f8

Observation c29f1b56-bd19-4960-bb1b-e28fac85f60b · outbound

This paper cites Marlin: Masked autoencoder for facial video representation learning,.

EXPOTION: Facial Expression and Motion Control for Multimodal Music Generation Marlin: Masked autoencoder for facial video representation learning,

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:40:09.043876Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T19:40:05.230502Z digest=sha256:8f67ab8b44844485b8c1bd77a68dc3d7e03c14c9bd0d6936ab94561d4adefde8

Observation 83659852-de37-4dcc-a2fc-0941b0e9eaea · outbound

This paper cites Synch- former: Efficient synchronization from sparse cues,.

EXPOTION: Facial Expression and Motion Control for Multimodal Music Generation Synch- former: Efficient synchronization from sparse cues,

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:40:08.910894Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T19:40:05.317236Z digest=sha256:b4ac898e1ec411fde51801d3075aaa3ccfacd435594061f139b76e5475f7bb97

Observation 8ee1745e-fb97-48c2-8d7c-8b64bbcdc768 · outbound

This paper cites Raft: Recurrent all-pairs field transforms for optical flow,.

EXPOTION: Facial Expression and Motion Control for Multimodal Music Generation Raft: Recurrent all-pairs field transforms for optical flow,

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:40:08.790842Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T19:40:05.387260Z digest=sha256:1137edb18d406e03b5bbf1e3e0692b3cd2d17755fa9dd8f4bb6d17b7c28e3068

Observation 30df38f1-ba3e-4b42-9166-3b4bc2c8f2e6 · outbound

This paper cites Salmonn: Towards generic hear- ing abilities for large language models,.

EXPOTION: Facial Expression and Motion Control for Multimodal Music Generation Salmonn: Towards generic hear- ing abilities for large language models,

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:40:08.638745Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T19:40:05.475020Z digest=sha256:c2c26b77b619c419ead887930098daff6b677d5ead4d67c27f617d5b6de1b83e

Observation a906a825-ceb7-4767-a58c-469be428b1c0 · outbound

This paper cites High Fidelity Neural Audio Compression.

EXPOTION: Facial Expression and Motion Control for Multimodal Music Generation High Fidelity Neural Audio Compression

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-06T19:40:05.562715Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:40:05.562715Z digest=sha256:10ca9380564b157089f2b9f873fbffd816335dbd00645c38a3c5ca891c2760a6

Observation 3570c2dd-fc12-4839-94ba-a278eac465da · outbound

This paper cites Video2music: Suitable music generation from videos using an affective multimodal transformer model,.

EXPOTION: Facial Expression and Motion Control for Multimodal Music Generation Video2music: Suitable music generation from videos using an affective multimodal transformer model,

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:40:08.478892Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T19:40:05.646574Z digest=sha256:50d951cf9b3e5b42adbe02d95476ad343babfaca7276954ee6ec0dadef396eaa

Observation 4d498afb-aabf-454d-b1f7-7cb6a779273d · outbound

This paper cites LanguageBind: Extending Video-Language Pretraining to N-modality by Language-based Semantic Alignment.

EXPOTION: Facial Expression and Motion Control for Multimodal Music Generation LanguageBind: Extending Video-Language Pretraining to N-modality by Language-based Semantic Alignment

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-06T19:40:06.423278Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:40:06.423278Z digest=sha256:7155bddf7a6a1f46ad07d18080a07950d3d94da9a4319822c5ec05a5a11f5a31

Observation 2109cd34-ce37-4fb4-ab09-158242398ce3 · outbound

This paper cites Frechet audio distance: A reference-free metric for evaluating music enhancement algorithms,.

EXPOTION: Facial Expression and Motion Control for Multimodal Music Generation Frechet audio distance: A reference-free metric for evaluating music enhancement algorithms,

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:40:08.324059Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T19:40:05.802192Z digest=sha256:c07f40ee1f135ae96769a39fed250d30d57d9824c9b5016081e9b126c454cf68

Observation b3b0825d-f5ef-4181-bdd0-f6dda41b6f6c · outbound

This paper cites Cnn architectures for large-scale audio classification,.

EXPOTION: Facial Expression and Motion Control for Multimodal Music Generation Cnn architectures for large-scale audio classification,

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:40:08.184873Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T19:40:05.915584Z digest=sha256:0808cef378ea0919740203cee3524897f9d46469318ce76ea2a14e70142c2afe

Observation 9bc0bd82-ac09-487d-ae91-6404e14bdf0a · outbound

This paper cites Panns: Large-scale pretrained audio neural networks for audio pattern recognition,.

EXPOTION: Facial Expression and Motion Control for Multimodal Music Generation Panns: Large-scale pretrained audio neural networks for audio pattern recognition,

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:40:08.043067Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T19:40:06.008469Z digest=sha256:c986c51afc8640a030f1c57bab33ddaf19c2c973dfe6bff6200692013412c82d

Observation 4ae24de1-8228-4f7d-ad82-b95dbfabf1d2 · outbound

This paper cites Efficient training of audio transform- ers with patchout,.

EXPOTION: Facial Expression and Motion Control for Multimodal Music Generation Efficient training of audio transform- ers with patchout,

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:40:07.890135Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T19:40:06.095672Z digest=sha256:5e90cbaa97cad6bb9e1e939422d00f32cce099a6ca630e568d4694dc55654277

Observation a7efb45e-843f-4d91-96e9-5a8b69baac36 · outbound

This paper cites Improved techniques for training gans,.

EXPOTION: Facial Expression and Motion Control for Multimodal Music Generation Improved techniques for training gans,

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:40:07.747301Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T19:40:06.183184Z digest=sha256:2b815f37cd4d09d9993f8aba93f86ae8334823302a5d52100a7fea3955b653dd

Observation ddce1c67-46ab-452f-9717-4bbfefbf307d · outbound

This paper cites Large-scale contrastive language-audio pretraining with feature fusion and keyword-to-caption augmentation,.

EXPOTION: Facial Expression and Motion Control for Multimodal Music Generation Large-scale contrastive language-audio pretraining with feature fusion and keyword-to-caption augmentation,

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:40:07.534298Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T19:40:06.270020Z digest=sha256:d0748ae0914cde7758e56079335d626dc5220f428c22995c92be337c00cb4954

Observation f793ccb1-7e6e-4408-82c8-6d38868b0d67 · outbound

This paper cites Available: https://arxiv.org/abs/2211.

EXPOTION: Facial Expression and Motion Control for Multimodal Music Generation Available: https://arxiv.org/abs/2211

Reference 2023

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:40:07.382378Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T19:40:06.341162Z digest=sha256:cc5cd612d6586ba0426a5c6b6e70bce9dd2a01dcd0f1081804c312e70252c12a

Observation 7a790e88-2b6f-4552-95e8-4e401d7c3a91 · outbound

This paper cites Available: http://dx.doi.org/10.1016/j.

EXPOTION: Facial Expression and Motion Control for Multimodal Music Generation Available: http://dx.doi.org/10.1016/j

Reference 2024

Resolution
malformed identifier
no resolver link, observed 2026-08-06T19:40:05.719311Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:40:05.719311Z digest=sha256:ce65a3af4975358c907f8c754b71e5a0bb802d6d6c2f533c1a1fa56b29fb59cc

Pith citing papers

Observation 9e40e093-95c3-427f-b276-9e5220945e43 · inbound

EXPOTION: Facial Expression and Motion Control for Multimodal Music Generation cites this paper.

EXPOTION: Facial Expression and Motion Control for Multimodal Music Generation EXPOTION: Facial Expression and Motion Control for Multimodal Music Generation

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-08-06T19:40:07.236634Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T19:40:02.091028Z digest=sha256:5fe1c0dd48636462e9b5ad5754dec57c13d102b054e81588de26f0363f9dcbd5