Pith. sign in

Paper Citation Record · LEDGER

MMAudio: Taming Multimodal Joint Training for High-Quality Video-to-Audio Synthesis

As of 10 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 11 inbound Pith citation observations for arXiv:2412.15322.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.15322 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 11 of 11 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 11 of 11 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-08T14:26:46.641742Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-22T23:07:14.501348Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 21929990-4d8e-41c6-992e-2b69a4e41bde · inbound

Lumina-Video: Efficient and Flexible Video Generation with Multi-scale Next-DiT cites this paper.

Lumina-Video: Efficient and Flexible Video Generation with Multi-scale Next-DiT MMAudio: Taming Multimodal Joint Training for High-Quality Video-to-Audio Synthesis

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-08T14:26:46.641742Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T14:26:46.641742Z digest=sha256:87458d1a7505e05a8df14b9cff5379ef1d75485b893dfea055fef20f210e9e43

Observation 91795adc-1b3d-410e-8887-008a4e0d6bb9 · inbound

Wan: Open and Advanced Large-Scale Video Generative Models cites this paper.

Wan: Open and Advanced Large-Scale Video Generative Models MMAudio: Taming Multimodal Joint Training for High-Quality Video-to-Audio Synthesis

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-22T23:07:14.505181Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-22T23:05:32.595632Z digest=sha256:b935d60e4c6285159da79f31deaeb243f91399e1ddd93a6a9da89392a85670e1

Observation 4ba353cf-9ed8-4907-825a-6f6705aa76b4 · inbound

SpecMaskFoley: Steering Pretrained Spectral Masked Generative Transformer Toward Synchronized Video-to-audio Synthesis via ControlNet cites this paper.

SpecMaskFoley: Steering Pretrained Spectral Masked Generative Transformer Toward Synchronized Video-to-audio Synthesis via ControlNet MMAudio: Taming Multimodal Joint Training for High-Quality Video-to-Audio Synthesis

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T15:09:54.999597Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:09:54.999597Z digest=sha256:2b098e8aa6158d02f7a8ad165d5ee2cda0f0705206457e4f54ee93b5272f1aeb

Observation 2309e642-f146-464a-a083-36fb2b903a66 · inbound

Towards Video to Piano Music Generation with Chain-of-Perform Support Benchmarks cites this paper.

Towards Video to Piano Music Generation with Chain-of-Perform Support Benchmarks MMAudio: Taming Multimodal Joint Training for High-Quality Video-to-Audio Synthesis

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T14:05:09.943750Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:05:09.943750Z digest=sha256:0f79e739e26d080fc42aa1aedfa24fc8c5562c0d58eee7d8951f9b41351fc27e

Observation 30f33d15-5961-462a-a3c1-2596da605e46 · inbound

AudioGenie: A Training-Free Multi-Agent Framework for Diverse Multimodality-to-Multiaudio Generation cites this paper.

AudioGenie: A Training-Free Multi-Agent Framework for Diverse Multimodality-to-Multiaudio Generation MMAudio: Taming Multimodal Joint Training for High-Quality Video-to-Audio Synthesis

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T13:21:53.035618Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:21:53.035618Z digest=sha256:c996df6852e648d4c96ac74b487e9f4b4869d3d104550527aca793f64544060c

Observation 14e806fc-87e0-48fa-9c31-79ec9b6b96c0 · inbound

Sounding that Object: Interactive Object-Aware Image to Audio Generation cites this paper.

Sounding that Object: Interactive Object-Aware Image to Audio Generation MMAudio: Taming Multimodal Joint Training for High-Quality Video-to-Audio Synthesis

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-07T10:53:14.895706Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:53:14.895706Z digest=sha256:4196bdcdcd1c47cf9c9395c3b62dc84a32dd654f5b57917ea81d8a7fa58b97ae

Observation 713fd661-c934-40de-b955-c4023f195efc · inbound

DualDub: Video-to-Soundtrack Generation via Joint Speech and Background Audio Synthesis cites this paper.

DualDub: Video-to-Soundtrack Generation via Joint Speech and Background Audio Synthesis MMAudio: Taming Multimodal Joint Training for High-Quality Video-to-Audio Synthesis

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T17:46:13.686519Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:46:13.686519Z digest=sha256:2410e86029db851a5dcdae6aa6b44a55f1dc4f50ac579dcb74f280663355d548

Observation 7244bc48-d4f4-41f9-9ec9-caa3a12af7b1 · inbound

FoleyDesigner: Immersive Stereo Foley Generation with Precise Spatio-Temporal Alignment for Film Clips cites this paper.

FoleyDesigner: Immersive Stereo Foley Generation with Precise Spatio-Temporal Alignment for Film Clips MMAudio: Taming Multimodal Joint Training for High-Quality Video-to-Audio Synthesis

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-10T23:40:51.764241Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T18:57:21.434793Z digest=sha256:136dffcfa0de008d03da4fad3a9fe3effc107537c2130878d85f067ab2a4a8a6

Observation 19f406e3-594d-419b-8900-c45e54a35f16 · inbound

TMD-Bench: A Multi-Level Evaluation Paradigm for Music-Dance Co-Generation cites this paper.

TMD-Bench: A Multi-Level Evaluation Paradigm for Music-Dance Co-Generation MMAudio: Taming Multimodal Joint Training for High-Quality Video-to-Audio Synthesis

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-11T16:36:06.093694Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-09T16:11:16.551910Z digest=sha256:2292ddbaba0f579d80f37a0b98bd09bae94c280d39bad46ef4ebc0ce0ae1eee9

Observation 85d98edd-c600-454c-a9dc-983ee4baa906 · inbound

SynSFX: Multi-Model Sound Effects Synthesis Dataset for Deepfake Detection and Evaluation cites this paper.

SynSFX: Multi-Model Sound Effects Synthesis Dataset for Deepfake Detection and Evaluation MMAudio: Taming Multimodal Joint Training for High-Quality Video-to-Audio Synthesis

Reference 22

Resolution
unresolved
no resolver link, observed 2026-07-11T12:34:20.057072Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T12:34:20.057072Z digest=sha256:3f81f6454f854c00e6457119c383cdfd9ab7ba11465f40d437ec31e3aa693ee4

Observation 8cfa0b25-621d-4ebf-a247-cdecb6287abc · inbound

KVAE: Family of Tokenizers for Multimodal Generative Models cites this paper.

KVAE: Family of Tokenizers for Multimodal Generative Models MMAudio: Taming Multimodal Joint Training for High-Quality Video-to-Audio Synthesis

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T23:24:48.957682Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:24:48.957682Z digest=sha256:65e659cdcb46c68c13ca5492875b2eb55e62d0e199d8f972f93b397b16d1b84b