Pith. sign in

Paper Citation Record · LEDGER

Mutual Forcing: Dual-Mode Self-Evolution for Fast Autoregressive Audio-Video Character Generation

As of 2 August 2026, this Paper Citation Record lists 61 of 61 outbound references and 1 inbound Pith citation observation for arXiv:2604.25819.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2604.25819 v1

Coverage vector

measured 61 of 61 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-07T16:54:23.108142Z

measured 62 of 62 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-02T06:30:47.504484+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-07-31T17:08:12.280455Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

61 of 61 outbound references displayed

  • verified exact35
  • verified fuzzy25
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 9483e7d1-ecde-4a64-9a8d-a154ed9e87c4 · outbound

This paper cites Seamless Interaction: Dyadic Audiovisual Motion Modeling and Large-Scale Dataset.

Mutual Forcing: Dual-Mode Self-Evolution for Fast Autoregressive Audio-Video Character Generation Seamless Interaction: Dyadic Audiovisual Motion Modeling and Large-Scale Dataset

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-11T23:31:14.745066Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-07T16:54:23.108142Z digest=sha256:c039883a7c887dfb101b775e0df2ab01b09c669cfda3b2918933603e20272ca4

Observation e4937179-044e-4778-9299-34a27c131827 · outbound

This paper cites FunAudioLLM: Voice Understanding and Generation Foundation Models for Natural Interaction Between Humans and LLMs.

Mutual Forcing: Dual-Mode Self-Evolution for Fast Autoregressive Audio-Video Character Generation FunAudioLLM: Voice Understanding and Generation Foundation Models for Natural Interaction Between Humans and LLMs

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-11T23:31:14.621422Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-07T16:54:23.108142Z digest=sha256:3be42eeef5df7f0165972f9cc5257a4b64e5029ff170e425436d10812de9f58b

Observation 63ea7ca0-06c5-4fa6-afe1-46caed50243d · outbound

This paper cites Genie: Generative interactive environments.

Mutual Forcing: Dual-Mode Self-Evolution for Fast Autoregressive Audio-Video Character Generation Genie: Generative interactive environments

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-05-27T01:38:40.451204Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-07T16:54:23.108142Z digest=sha256:43de28286e1b1458621dc135bb557fc75049c51d97c5c7a5b6fb980054fac1bb

Observation 11183744-dcdc-4caf-8842-9add4b4aa067 · outbound

This paper cites Diffusion forcing: Next-token prediction meets full-sequence diffusion.Advances in Neural Information Processing Systems, 37:24081–24125.

Mutual Forcing: Dual-Mode Self-Evolution for Fast Autoregressive Audio-Video Character Generation Diffusion forcing: Next-token prediction meets full-sequence diffusion.Advances in Neural Information Processing Systems, 37:24081–24125

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-05-27T01:38:40.426402Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-07T16:54:23.108142Z digest=sha256:4849283b3fefc27efdeb93f24145477632a423f2cd6873f708236088d3f33adf

Observation 5608cb4f-9f2c-4b1f-a238-2a3e90616bda · outbound

This paper cites SkyReels-V2: Infinite-length Film Generative Model.

Mutual Forcing: Dual-Mode Self-Evolution for Fast Autoregressive Audio-Video Character Generation SkyReels-V2: Infinite-length Film Generative Model

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-14T20:23:04.486750Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-07T16:54:23.108142Z digest=sha256:bde00351aa5d438304cf6dd166a03b8a396158e402c9152160319bb000007efc

Observation f10ab449-5950-4997-aa66-81ecd43ccdc3 · outbound

This paper cites Panda-70m: Captioning 70m videos with multiple cross- modality teachers.

Mutual Forcing: Dual-Mode Self-Evolution for Fast Autoregressive Audio-Video Character Generation Panda-70m: Captioning 70m videos with multiple cross- modality teachers

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-05-27T01:38:39.903640Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-07T16:54:23.108142Z digest=sha256:8b45cf6d263c2cc54778f5045e2940af8ed9ca0388e65b4c3a0acddd2fc056b0

Observation 06ccb3b0-3007-403b-b9e3-ccaca911d14d · outbound

This paper cites Out of time: automated lip sync in the wild.

Mutual Forcing: Dual-Mode Self-Evolution for Fast Autoregressive Audio-Video Character Generation Out of time: automated lip sync in the wild

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-05-27T01:38:39.907186Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-07T16:54:23.108142Z digest=sha256:ddcd75186e4bf8b9869df63b32835f79a498135262d1da4e8409a779f6f529b4

Observation 238f6a01-a0a2-4728-9bd7-0a7495a9d4c7 · outbound

This paper cites Self-Forcing++: Towards Minute-Scale High-Quality Video Generation.

Mutual Forcing: Dual-Mode Self-Evolution for Fast Autoregressive Audio-Video Character Generation Self-Forcing++: Towards Minute-Scale High-Quality Video Generation

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-15T22:39:54.578381Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-07T16:54:23.108142Z digest=sha256:f9b55ed5fd59bee7fb539df83d74b1df3c5ae6cac6e555a002053704672dc720

Observation 6c39ea89-43a9-4031-a060-bd2288889756 · outbound

This paper cites Stable audio open.

Mutual Forcing: Dual-Mode Self-Evolution for Fast Autoregressive Audio-Video Character Generation Stable audio open

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-05-27T01:38:39.887095Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-07T16:54:23.108142Z digest=sha256:3fd9d29d883ccdc8c7f9d4cb6e72ad2a26db15c43b3fcf8686be9c11e708daee

Observation 6b29502d-9a96-4714-a627-086ad4d3fd45 · outbound

This paper cites Phased dmd: Few-step distribution matching distillation via score matching within subintervals.arXiv preprint arXiv:2510.27684.

Mutual Forcing: Dual-Mode Self-Evolution for Fast Autoregressive Audio-Video Character Generation Phased dmd: Few-step distribution matching distillation via score matching within subintervals.arXiv preprint arXiv:2510.27684

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-11T23:31:14.730777Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-07T16:54:23.108142Z digest=sha256:25b97653ad82106dae895a1d12445fc544eabf19a09a1bf37ec0dcb9748477ae

Observation ca8f6667-f31d-48c6-ab2b-35604bb8e0b5 · outbound

This paper cites One Step Diffusion via Shortcut Models.

Mutual Forcing: Dual-Mode Self-Evolution for Fast Autoregressive Audio-Video Character Generation One Step Diffusion via Shortcut Models

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-15T06:40:55.602339Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-07T16:54:23.108142Z digest=sha256:f0dc18dde1b8305717a5957f807b8198675880dc6b1aed0cfe5ceb3916a4cdb1

Observation 5eed8259-7bb5-4bd3-8c7d-6fdfe4dc5f2e · outbound

This paper cites OmniAvatar: Efficient Audio-Driven Avatar Video Generation with Adaptive Body Animation.

Mutual Forcing: Dual-Mode Self-Evolution for Fast Autoregressive Audio-Video Character Generation OmniAvatar: Efficient Audio-Driven Avatar Video Generation with Adaptive Body Animation

Reference 12

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T23:31:15.014324Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-07T16:54:23.108142Z digest=sha256:16776ea1d6221068f585509b246b0ca5d1daa905d6b86b74f515c46667e6d1a6

Observation 3ea3c7b8-d52c-47e6-9c48-0973c96485da · outbound

This paper cites Wan-S2V: Audio-Driven Cinematic Video Generation.

Mutual Forcing: Dual-Mode Self-Evolution for Fast Autoregressive Audio-Video Character Generation Wan-S2V: Audio-Driven Cinematic Video Generation

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-11T23:31:14.635424Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-07T16:54:23.108142Z digest=sha256:80d67765177ad31ad28cad578d80a39bafde4cad53eca3f11c4ed9400bfb6fbf

Observation d93638e7-fb24-462a-87de-b99b513f5f56 · outbound

This paper cites Long video generation with time-agnostic vqgan and time-sensitive transformer.

Mutual Forcing: Dual-Mode Self-Evolution for Fast Autoregressive Audio-Video Character Generation Long video generation with time-agnostic vqgan and time-sensitive transformer

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-05-27T01:38:39.893390Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-07T16:54:23.108142Z digest=sha256:0aaf870b7fed489d366fa2e76c90d33eb4d2de84eb180fcabc6ac6e3a3bd1afd

Observation e4dc32b5-e10e-4ded-bf6f-f373898f5324 · outbound

This paper cites Sparsectrl: Adding sparse controls to text-to-video diffusion models.

Mutual Forcing: Dual-Mode Self-Evolution for Fast Autoregressive Audio-Video Character Generation Sparsectrl: Adding sparse controls to text-to-video diffusion models

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-05-27T01:38:40.415545Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-07T16:54:23.108142Z digest=sha256:3cc1c28be149317a69b6f6686b8a35a61318c34b9318ee183a4e9674f26e07b2

Observation 0520f6c2-adbd-4266-a494-89d44b5f757e · outbound

This paper cites AnimateDiff: Animate Your Personalized Text-to-Image Diffusion Models without Specific Tuning.

Mutual Forcing: Dual-Mode Self-Evolution for Fast Autoregressive Audio-Video Character Generation AnimateDiff: Animate Your Personalized Text-to-Image Diffusion Models without Specific Tuning

Reference 16

Resolution
verified exact
local_arxiv, observed 2026-05-11T23:31:14.843543Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-07T16:54:23.108142Z digest=sha256:af0d458aaaf63fc651d6c8a5f010b5a6f77ce65bf361a1d279b5974547555b68

Observation 83c3caef-ab5c-4c84-b7f1-2c300b24aaf0 · outbound

This paper cites Av-link: Temporally-aligned diffusion features for cross-modal audio-video generation.

Mutual Forcing: Dual-Mode Self-Evolution for Fast Autoregressive Audio-Video Character Generation Av-link: Temporally-aligned diffusion features for cross-modal audio-video generation

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-05-27T01:38:40.422504Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-07T16:54:23.108142Z digest=sha256:d102150605a032fce90f11d24e7dd9edf3f9acc3634d2a0aafe7eef3036d25a9

Observation 98a472dd-2c41-4323-a75e-24c8481337a1 · outbound

This paper cites Emilia: An extensive, multilingual, and diverse speech dataset for large-scale speech generation.

Mutual Forcing: Dual-Mode Self-Evolution for Fast Autoregressive Audio-Video Character Generation Emilia: An extensive, multilingual, and diverse speech dataset for large-scale speech generation

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-05-27T01:38:40.430814Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-07T16:54:23.108142Z digest=sha256:f253598858b9c0dfdf11adcd798ac3940a21990b49d376a0a2fd3aea775cf45c

Observation a4343819-02f6-4907-845f-a29e5d72e5b8 · outbound

This paper cites Latent Video Diffusion Models for High-Fidelity Long Video Generation.

Mutual Forcing: Dual-Mode Self-Evolution for Fast Autoregressive Audio-Video Character Generation Latent Video Diffusion Models for High-Fidelity Long Video Generation

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-15T04:27:43.523852Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-07T16:54:23.108142Z digest=sha256:ba4dce13b6e9ce726120e253d0e0ccf902a51ebb01b775e5e3b56dae639afd7a

Observation f8a3f2c4-b729-4546-a552-dfb9175342ed · outbound

This paper cites Imagen Video: High Definition Video Generation with Diffusion Models.

Mutual Forcing: Dual-Mode Self-Evolution for Fast Autoregressive Audio-Video Character Generation Imagen Video: High Definition Video Generation with Diffusion Models

Reference 20

Resolution
verified exact
local_arxiv, observed 2026-05-11T23:31:14.865612Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-07T16:54:23.108142Z digest=sha256:708806c37fa31f1532962db9993d0b524b35707980d85fea911c9d25646c3fe0

Observation 964cdf16-eae7-473d-a1b3-ab24aa52dba7 · outbound

This paper cites Video diffusion models.Advances in neural information processing systems, 35:8633–8646.

Mutual Forcing: Dual-Mode Self-Evolution for Fast Autoregressive Audio-Video Character Generation Video diffusion models.Advances in neural information processing systems, 35:8633–8646

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-05-27T01:38:40.437542Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-07T16:54:23.108142Z digest=sha256:b164126d88ae602975c07b8632427b85a16b14954521766e466eb852131bcce4

Observation 1e418354-9b7f-4dcd-9b70-a4514e318eb1 · outbound

This paper cites CogVideo: Large-scale Pretraining for Text-to-Video Generation via Transformers.

Mutual Forcing: Dual-Mode Self-Evolution for Fast Autoregressive Audio-Video Character Generation CogVideo: Large-scale Pretraining for Text-to-Video Generation via Transformers

Reference 22

Resolution
verified exact
local_arxiv, observed 2026-05-11T23:31:14.770777Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-07T16:54:23.108142Z digest=sha256:6c5816562ca095e8ee9702902450056367a53a2cf55860ffb03935d72e28a590

Observation ee8c4a9c-1b08-4428-9553-2b28688cab89 · outbound

This paper cites Self Forcing: Bridging the Train-Test Gap in Autoregressive Video Diffusion.

Mutual Forcing: Dual-Mode Self-Evolution for Fast Autoregressive Audio-Video Character Generation Self Forcing: Bridging the Train-Test Gap in Autoregressive Video Diffusion

Reference 23

Resolution
verified exact
local_arxiv, observed 2026-05-11T23:31:14.874434Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-07T16:54:23.108142Z digest=sha256:5f507747962fcaf4a463bc4de11bbb78b06842a8792e8b6eaa006bdf0b0cdac1

Observation 19ec8ee7-9b08-4f8c-8865-d6376b661e51 · outbound

This paper cites Jova: Unified multimodal learning for joint video-audio generation.

Mutual Forcing: Dual-Mode Self-Evolution for Fast Autoregressive Audio-Video Character Generation Jova: Unified multimodal learning for joint video-audio generation

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-11T23:31:14.838308Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-07T16:54:23.108142Z digest=sha256:1442622977780196cc5bff7f864517c2c24fcde82fe00ac21ff3c81aae49a628

Observation eaa8a407-2381-4b19-b068-3a413431877b · outbound

This paper cites Vbench: Comprehensive benchmark suite for video generative models.

Mutual Forcing: Dual-Mode Self-Evolution for Fast Autoregressive Audio-Video Character Generation Vbench: Comprehensive benchmark suite for video generative models

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-05-27T01:38:40.440887Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-07T16:54:23.108142Z digest=sha256:72eafce6d57b9ee62e38bdcb918a32a77abe88a8ba07b8b2a0dc84c2dafb277c

Observation 760f805d-15ad-4f8a-b7cf-a5bd941f6441 · outbound

This paper cites A simple but strong baseline for sounding video generation: Effective adaptation of audio and video diffusion models for joint generation.

Mutual Forcing: Dual-Mode Self-Evolution for Fast Autoregressive Audio-Video Character Generation A simple but strong baseline for sounding video generation: Effective adaptation of audio and video diffusion models for joint generation

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-05-27T01:38:39.900309Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-07T16:54:23.108142Z digest=sha256:eb1896be445c6d8693fac5e1a2614a5de43b83a8beec3dc9e869497d5786264a

Observation c957a1f6-3084-485a-93d8-fca6671176dc · outbound

This paper cites Pyramidal flow matching for efficient video generative modeling.arXiv preprint arXiv:2410.05954.

Mutual Forcing: Dual-Mode Self-Evolution for Fast Autoregressive Audio-Video Character Generation Pyramidal flow matching for efficient video generative modeling.arXiv preprint arXiv:2410.05954

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-11T23:31:14.778880Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-07T16:54:23.108142Z digest=sha256:8c4e4f2b4b5b3552f2124cc01ff2c1f6e8b1b7c065fdd0f2046fd5d9254008b8

Observation 09789276-601d-42c6-8be1-c492bd2449e0 · outbound

This paper cites HunyuanVideo: A Systematic Framework For Large Video Generative Models.

Mutual Forcing: Dual-Mode Self-Evolution for Fast Autoregressive Audio-Video Character Generation HunyuanVideo: A Systematic Framework For Large Video Generative Models

Reference 28

Resolution
verified exact
local_arxiv, observed 2026-05-11T23:31:14.899882Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-07T16:54:23.108142Z digest=sha256:47d148a508727b4dd2cf9840750b8c16c21c1c583dfa49e82832a42cf883afbf

Observation 7bc17ca6-66cc-4c87-90ca-818a8bd4a5c6 · outbound

This paper cites JavisDiT: Joint audio-video diffusion transformer with hierarchical spatio-temporal prior synchronization.

Mutual Forcing: Dual-Mode Self-Evolution for Fast Autoregressive Audio-Video Character Generation JavisDiT: Joint audio-video diffusion transformer with hierarchical spatio-temporal prior synchronization

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-11T23:31:14.798344Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-07T16:54:23.108142Z digest=sha256:54a8f83f3c36ce1b06935bf011e9dd96f2f2d26ecc9e88425bdda503623da1ca

Observation b055e0a9-b452-43ce-81c5-d471e19ec570 · outbound

This paper cites Ovi: Twin Backbone Cross-Modal Fusion for Audio-Video Generation.

Mutual Forcing: Dual-Mode Self-Evolution for Fast Autoregressive Audio-Video Character Generation Ovi: Twin Backbone Cross-Modal Fusion for Audio-Video Generation

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-05-17T01:03:15.294735Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-07T16:54:23.108142Z digest=sha256:cdfcec9140d393190bf821e7814c97306aff3a3e2547f80c7b038ed3d75fb37f

Observation 573c70d1-a835-40ff-8d43-0df409fe0d2a · outbound

This paper cites Follow your pose: Pose- guided text-to-video generation using pose-free videos.

Mutual Forcing: Dual-Mode Self-Evolution for Fast Autoregressive Audio-Video Character Generation Follow your pose: Pose- guided text-to-video generation using pose-free videos

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-05-27T01:38:40.457324Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-07T16:54:23.108142Z digest=sha256:2636bec8f303308776293e0df043a85768d4a49ce6c0af2ad0f9eccde290e5e4

Observation 9781c4f6-b08d-41ea-ac15-add02e3e05e8 · outbound

This paper cites Mofa-video: Controllable image animation via generative motion field adaptions in frozen image-to-video diffusion model.

Mutual Forcing: Dual-Mode Self-Evolution for Fast Autoregressive Audio-Video Character Generation Mofa-video: Controllable image animation via generative motion field adaptions in frozen image-to-video diffusion model

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-05-27T01:38:39.871011Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-07T16:54:23.108142Z digest=sha256:d4e2e66e006f147dc7df3a1c0b4caa3eeff10b9459c7f1e72feeb63be332c61e

Observation 612cc8b5-273e-4308-b256-4a82a96e9a4f · outbound

This paper cites Video generation models as world simulators.

Mutual Forcing: Dual-Mode Self-Evolution for Fast Autoregressive Audio-Video Character Generation Video generation models as world simulators

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-05-27T01:38:39.889097Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-07T16:54:23.108142Z digest=sha256:9a96349b1dfa50240abc1199f1bbb2254c90d8a912ae7c803b0109bcb77d3e14

Observation a2386faf-cc56-4f20-9ee5-c983048e2f48 · outbound

This paper cites Robust Speech Recognition via Large-Scale Weak Supervision.

Mutual Forcing: Dual-Mode Self-Evolution for Fast Autoregressive Audio-Video Character Generation Robust Speech Recognition via Large-Scale Weak Supervision

Reference 34

Resolution
verified exact
local_arxiv, observed 2026-05-11T23:31:14.647350Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-07T16:54:23.108142Z digest=sha256:9146705029c327110e7602a843df33502bc0e00ef2f6582a5185487b8a349de4

Observation 63b3ad56-8949-4c4a-b915-0d2cd404cc00 · outbound

This paper cites ConsistI2V: Enhancing Visual Consistency for Image-to-Video Generation.

Mutual Forcing: Dual-Mode Self-Evolution for Fast Autoregressive Audio-Video Character Generation ConsistI2V: Enhancing Visual Consistency for Image-to-Video Generation

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-05-11T23:31:14.829831Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-07T16:54:23.108142Z digest=sha256:609f2e2653d7c789f49dbc3b8cd24a53e4cd847bf0675447c69e72f7720a25a3

Observation 0119bd0b-973e-4771-8912-fa9104d7ff9d · outbound

This paper cites High-resolution image synthesis with latent diffusion models.

Mutual Forcing: Dual-Mode Self-Evolution for Fast Autoregressive Audio-Video Character Generation High-resolution image synthesis with latent diffusion models

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-05-27T01:38:40.434167Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-07T16:54:23.108142Z digest=sha256:bbe112e5d089be5f039f072733055e106c26c472961c6affa9ed4015f4e83b46

Observation 06cf7f12-db15-4046-8f60-20feb2c747b7 · outbound

This paper cites Mm-diffusion: Learning multi-modal diffusion models for joint audio and video generation.

Mutual Forcing: Dual-Mode Self-Evolution for Fast Autoregressive Audio-Video Character Generation Mm-diffusion: Learning multi-modal diffusion models for joint audio and video generation

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-05-27T01:38:39.882387Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-07T16:54:23.108142Z digest=sha256:7f4b9c036e46ebea4d8344687b0bb4130c30df4e8d9f9723e176b7b6a139f8b3

Observation 70be4f20-b11b-420a-bdcf-54e47d56337b · outbound

This paper cites Improved Techniques for Training Consistency Models.

Mutual Forcing: Dual-Mode Self-Evolution for Fast Autoregressive Audio-Video Character Generation Improved Techniques for Training Consistency Models

Reference 38

Resolution
verified exact
arxiv_id, observed 2026-05-21T05:04:10.684842Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-07T16:54:23.108142Z digest=sha256:5a2f2649f9101da87a597674b574c642920b888989a02dd2ec65d822d5a7eca0

Observation 491aa484-adbe-4ae4-9753-e32dd20ddc57 · outbound

This paper cites RoFormer: Enhanced Transformer with Rotary Position Embedding.

Mutual Forcing: Dual-Mode Self-Evolution for Fast Autoregressive Audio-Video Character Generation RoFormer: Enhanced Transformer with Rotary Position Embedding

Reference 39

Resolution
verified exact
local_arxiv, observed 2026-05-11T23:31:14.708350Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-07T16:54:23.108142Z digest=sha256:1e0c9dac56828230ab609d56a73b1a82fd841fe4c1c5420653ea4e29a1cb2fc8

Observation 49efc2da-e127-4b97-8ca8-7b7eac439b59 · outbound

This paper cites MAGI-1: Autoregressive Video Generation at Scale.

Mutual Forcing: Dual-Mode Self-Evolution for Fast Autoregressive Audio-Video Character Generation MAGI-1: Autoregressive Video Generation at Scale

Reference 40

Resolution
verified exact
arxiv_id, observed 2026-05-13T20:31:15.871842Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-07T16:54:23.108142Z digest=sha256:5623200b2dfa93b42a83448d7efdd95fc879b2eadff69f2304e75ae579fd47bf

Observation b51213cc-ba48-4d17-b898-6bb0878f3b66 · outbound

This paper cites Meta audiobox aesthetics: Unified automatic quality assessment for speech, music, and sound.

Mutual Forcing: Dual-Mode Self-Evolution for Fast Autoregressive Audio-Video Character Generation Meta audiobox aesthetics: Unified automatic quality assessment for speech, music, and sound

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-05-27T01:38:39.896690Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-07T16:54:23.108142Z digest=sha256:963b49bda42d6652ba71abe5b015f5c68b87cbb59bda133531189c41f0044a77

Observation 8ae6c1b2-3d2c-43f0-afac-4ba7c2a26f30 · outbound

This paper cites Generating the future with adversarial transformers.

Mutual Forcing: Dual-Mode Self-Evolution for Fast Autoregressive Audio-Video Character Generation Generating the future with adversarial transformers

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-05-27T01:38:39.874789Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-07T16:54:23.108142Z digest=sha256:88cda7090dd87affbeb6b9f6063ce845d134022b2e2a3b52e7c7381490743f79

Observation 0d5194b7-1cd7-4332-a767-34932e0d06d5 · outbound

This paper cites Wan: Open and Advanced Large-Scale Video Generative Models.

Mutual Forcing: Dual-Mode Self-Evolution for Fast Autoregressive Audio-Video Character Generation Wan: Open and Advanced Large-Scale Video Generative Models

Reference 43

Resolution
verified exact
local_arxiv, observed 2026-05-11T23:31:14.887388Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-07T16:54:23.108142Z digest=sha256:71ef8ae22c54919ba0c193f3cb4a367cbc15712933cf6817e2eb746e887dd15a

Observation 408d256f-81e6-4aba-af42-0522fcafa1ac · outbound

This paper cites UniVerse-1: Unified Audio-Video Generation via Stitching of Experts.

Mutual Forcing: Dual-Mode Self-Evolution for Fast Autoregressive Audio-Video Character Generation UniVerse-1: Unified Audio-Video Generation via Stitching of Experts

Reference 44

Resolution
verified exact
arxiv_id, observed 2026-05-11T23:31:14.791158Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-07T16:54:23.108142Z digest=sha256:0e371a45a0ea574a369232ec877a7e5a88d696629b97247b026257e7a29e546a

Observation ee1021be-e022-4597-b536-0894dab58480 · outbound

This paper cites ModelScope Text-to-Video Technical Report.

Mutual Forcing: Dual-Mode Self-Evolution for Fast Autoregressive Audio-Video Character Generation ModelScope Text-to-Video Technical Report

Reference 45

Resolution
verified exact
arxiv_id, observed 2026-05-12T19:47:29.701736Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-07T16:54:23.108142Z digest=sha256:14255af0bdf84a59f59d9ec231eaf1d789762a003a9bc04d347b66ee9e77b6be

Observation fccbb80b-9232-48b2-87ea-b302e0750ebe · outbound

This paper cites Av-dit: Taming image diffusion transformers for efficient joint audio and video generation.

Mutual Forcing: Dual-Mode Self-Evolution for Fast Autoregressive Audio-Video Character Generation Av-dit: Taming image diffusion transformers for efficient joint audio and video generation

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-05-27T01:38:40.411736Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-07T16:54:23.108142Z digest=sha256:a840f650269455911f3f863450e6bbaa1104ceee16c46dd6b6b4f75e038d2a51

Observation e2398dca-1a8a-4c2e-87bd-5b4b26d2db5d · outbound

This paper cites Fantasytalk- ing: Realistic talking portrait generation via coherent motion synthesis.

Mutual Forcing: Dual-Mode Self-Evolution for Fast Autoregressive Audio-Video Character Generation Fantasytalk- ing: Realistic talking portrait generation via coherent motion synthesis

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-05-27T01:38:39.885831Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-07T16:54:23.108142Z digest=sha256:2e03f336e7a97e7eb41d93f72826bae6fdf98da852f5614a72968ad65d48a6bf

Observation 696ea226-a470-44a2-93e5-ff37edcbae03 · outbound

This paper cites Emu3: Next-Token Prediction is All You Need.

Mutual Forcing: Dual-Mode Self-Evolution for Fast Autoregressive Audio-Video Character Generation Emu3: Next-Token Prediction is All You Need

Reference 48

Resolution
verified exact
local_arxiv, observed 2026-05-11T23:31:15.005455Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-07T16:54:23.108142Z digest=sha256:cc64616f037fc2511b5271820ad37d013257eae4d1284c6d849ac5df37070357

Observation 3f6d8964-8e9a-444a-af15-d1d4f59abbff · outbound

This paper cites InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation.

Mutual Forcing: Dual-Mode Self-Evolution for Fast Autoregressive Audio-Video Character Generation InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation

Reference 49

Resolution
verified exact
arxiv_id, observed 2026-05-15T06:30:22.925421Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-07T16:54:23.108142Z digest=sha256:c3880a2a1e25f388fdf558e9163897c07b464dbfabc0c846c197825a87a32452

Observation abc3491d-9480-4f38-b430-594dc4f6170b · outbound

This paper cites Loong: Generating Minute-level Long Videos with Autoregressive Language Models.

Mutual Forcing: Dual-Mode Self-Evolution for Fast Autoregressive Audio-Video Character Generation Loong: Generating Minute-level Long Videos with Autoregressive Language Models

Reference 50

Resolution
verified exact
arxiv_id, observed 2026-05-11T23:31:14.905365Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-07T16:54:23.108142Z digest=sha256:0a673263cc65828bcc00debca7436b0427d7610f9ca361e937e67a3c442c0625

Observation 8be701ff-7d51-4c2f-b2ad-f23c0915b662 · outbound

This paper cites Large-scale contrastive language-audio pretraining with feature fusion and keyword-to-caption augmentation.

Mutual Forcing: Dual-Mode Self-Evolution for Fast Autoregressive Audio-Video Character Generation Large-scale contrastive language-audio pretraining with feature fusion and keyword-to-caption augmentation

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-05-27T01:38:40.447676Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-07T16:54:23.108142Z digest=sha256:930bb0cd356103b537fa66dae496db9ff215f24c2694953e8e3a3a49e1a1e27a

Observation ca8250ac-ccaa-4c88-b3e0-c958293f50c5 · outbound

This paper cites Magicanimate: Temporally consistent human image animation using diffusion model.

Mutual Forcing: Dual-Mode Self-Evolution for Fast Autoregressive Audio-Video Character Generation Magicanimate: Temporally consistent human image animation using diffusion model

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-05-27T01:38:40.418902Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-07T16:54:23.108142Z digest=sha256:fdf7769925d71190fffc4a8ab30166ba9c6f88b02ef31a6da90c1816dbe29110

Observation 74f02384-402b-4bf5-95ab-b4363f85e979 · outbound

This paper cites Stand-in: A lightweight and plug-and-play identity control for video generation.

Mutual Forcing: Dual-Mode Self-Evolution for Fast Autoregressive Audio-Video Character Generation Stand-in: A lightweight and plug-and-play identity control for video generation

Reference 53

Resolution
verified exact
arxiv_id, observed 2026-05-11T23:31:14.824208Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-07T16:54:23.108142Z digest=sha256:af3977e8c89138273586d3e427aee0edc33c61d081495b7f55af3823d8a32c9d

Observation ddc478f6-5966-456e-bf41-2a4c9617bcad · outbound

This paper cites LongLive: Real-time Interactive Long Video Generation.

Mutual Forcing: Dual-Mode Self-Evolution for Fast Autoregressive Audio-Video Character Generation LongLive: Real-time Interactive Long Video Generation

Reference 54

Resolution
verified exact
arxiv_id, observed 2026-05-15T03:52:59.837830Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-07T16:54:23.108142Z digest=sha256:fdadac7aeab9b812dd04275d3bbbe0e5de910ad4580cc73e8dcb3ee9f960de7f

Observation de5d3010-744d-4255-99d4-d9d60b850901 · outbound

This paper cites CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer.

Mutual Forcing: Dual-Mode Self-Evolution for Fast Autoregressive Audio-Video Character Generation CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer

Reference 55

Resolution
verified exact
local_arxiv, observed 2026-05-11T23:31:14.895425Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-07T16:54:23.108142Z digest=sha256:eca331f0522f7bf36ed1fa00ba91ce13e86a5c5add7e2cadb6143f982d3009ab

Observation 6b1ccfaa-9ab3-4e21-9095-c5a33dcbcbb0 · outbound

This paper cites Improved distribution matching distillation for fast image synthesis.Advances in neural information processing systems, 37:47455– 47487.

Mutual Forcing: Dual-Mode Self-Evolution for Fast Autoregressive Audio-Video Character Generation Improved distribution matching distillation for fast image synthesis.Advances in neural information processing systems, 37:47455– 47487

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-05-27T01:38:40.454525Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-07T16:54:23.108142Z digest=sha256:a514170d8a275a2f8ee8ebef5417025206355e463ea76e315b3c74e5bdb37738

Observation b76796d0-8582-4c25-8f6e-480186ddfdc0 · outbound

This paper cites From slow bidirectional to fast autoregressive video diffusion models.

Mutual Forcing: Dual-Mode Self-Evolution for Fast Autoregressive Audio-Video Character Generation From slow bidirectional to fast autoregressive video diffusion models

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-05-27T01:38:40.444157Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-07T16:54:23.108142Z digest=sha256:518876d2f73ef56bd1726e1ce9bb2849d5e2c71037a2830ba8d0d3d5b1c5b837

Observation c0df8de2-616e-4987-820f-dbaaae88734c · outbound

This paper cites arXiv preprint arXiv:2511.03334 (2025).

Mutual Forcing: Dual-Mode Self-Evolution for Fast Autoregressive Audio-Video Character Generation arXiv preprint arXiv:2511.03334 (2025)

Reference 58

Resolution
verified exact
arxiv_id, observed 2026-05-11T23:31:14.612127Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-07T16:54:23.108142Z digest=sha256:d2676e5a86386feffb435d06863969375dc9dfc43ef10eef097075b87b3741b1

Observation dc282b02-29ad-4377-8da4-1c1574feff64 · outbound

This paper cites SpeakerVid-5M: A Large-Scale High-Quality Dataset for Audio-Visual Dyadic Interactive Human Generation.

Mutual Forcing: Dual-Mode Self-Evolution for Fast Autoregressive Audio-Video Character Generation SpeakerVid-5M: A Large-Scale High-Quality Dataset for Audio-Visual Dyadic Interactive Human Generation

Reference 59

Resolution
verified exact
arxiv_id, observed 2026-05-11T23:31:14.592558Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-07T16:54:23.108142Z digest=sha256:f0882a8bd6cf081398b1a337d59fe93cf5461fb0a9dd6470aeda55d4ad386311

Observation dc8f86c8-0bc9-48c0-8db0-1b8db3b36212 · outbound

This paper cites Large Scale Diffusion Distillation via Score-Regularized Continuous-Time Consistency.

Mutual Forcing: Dual-Mode Self-Evolution for Fast Autoregressive Audio-Video Character Generation Large Scale Diffusion Distillation via Score-Regularized Continuous-Time Consistency

Reference 60

Resolution
verified exact
local_arxiv, observed 2026-05-11T23:31:14.675668Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-07T16:54:23.108142Z digest=sha256:2275e3a725d330b78f33a7c252b7a5448e8d9cb7e2c1df402de21add1b2caa05

Observation 93bb69bc-7543-435c-a1b8-111094802603 · outbound

This paper cites Allegro: Open the Black Box of Commercial-Level Video Generation Model.

Mutual Forcing: Dual-Mode Self-Evolution for Fast Autoregressive Audio-Video Character Generation Allegro: Open the Black Box of Commercial-Level Video Generation Model

Reference 61

Resolution
verified exact
arxiv_id, observed 2026-05-11T23:31:14.696999Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-07T16:54:23.108142Z digest=sha256:27831cbb8d6013f50af0855f1d140d47c69ebe21c6568456e1c3c2b9c5a8816a

Pith citing papers

Observation c5b24e07-50b4-45b5-83b0-022cd6e8aa5b · inbound

TaoMate: Anchor-Guided Memory Bridging Evolving and Reference States for Real-Time Audio-Video Digital Human Generation cites this paper.

TaoMate: Anchor-Guided Memory Bridging Evolving and Reference States for Real-Time Audio-Video Digital Human Generation Mutual Forcing: Dual-Mode Self-Evolution for Fast Autoregressive Audio-Video Character Generation

Reference 18

Resolution
unresolved
no resolver link, observed 2026-07-31T17:08:12.280455Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-31T17:08:12.280455Z digest=sha256:279ef435cd74fb9003cb59396d1b08d3e40a1b7f09497bc0109aff1e55983101