Pith. sign in

Paper Citation Record · LEDGER

Improving Video Diffusion Transformer Training by Multi-Feature Fusion and Alignment from Self-Supervised Vision Encoders

As of 22 August 2026, this Paper Citation Record lists 38 of 38 outbound references and 0 inbound Pith citation observations for arXiv:2509.09547.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2509.09547 v1

Coverage vector

measured 38 of 38 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-04T18:58:11.054269Z

measured 38 of 38 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

38 of 38 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved38
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 26eb7b67-5785-468d-ba32-3b6f849bf9e8 · outbound

This paper cites Goku: Flow Based Video Generative Foundation Models.

Improving Video Diffusion Transformer Training by Multi-Feature Fusion and Alignment from Self-Supervised Vision Encoders Goku: Flow Based Video Generative Foundation Models

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-04T18:58:10.893526Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:58:10.893526Z digest=sha256:e2f7fd7283067f5ba6fdb3cb63bf83231cdde248a370367b217b246ab06de450

Observation 0ed44e74-360b-4127-8c72-81f4380afe1a · outbound

This paper cites Motivated by this trend, we also analyzed the vision encoder used in a state-of-the-art MLLM to better understand its po- tential for video generation tasks.

Improving Video Diffusion Transformer Training by Multi-Feature Fusion and Alignment from Self-Supervised Vision Encoders Motivated by this trend, we also analyzed the vision encoder used in a state-of-the-art MLLM to better understand its po- tential for video generation tasks

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-04T18:58:11.054269Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:58:11.054269Z digest=sha256:127e9ebbfd20f67d8156be5c302a2d70bc275d8ec16238b0f3027f36547ba0ba

Observation 7a6f59ae-cc4d-443a-b019-bf176068a81e · outbound

This paper cites TokenFlow: Consistent Diffusion Features for Consistent Video Editing.

Improving Video Diffusion Transformer Training by Multi-Feature Fusion and Alignment from Self-Supervised Vision Encoders TokenFlow: Consistent Diffusion Features for Consistent Video Editing

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-04T18:58:10.903298Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:58:10.903298Z digest=sha256:934a75dec4553faca8be7ee673d2699379caf6041640832c8c8e4a393b2649f2

Observation c02f7f1d-9315-4a20-8b56-bd70303c5a9e · outbound

This paper cites AnimateDiff: Animate Your Personalized Text-to-Image Diffusion Models without Specific Tuning.

Improving Video Diffusion Transformer Training by Multi-Feature Fusion and Alignment from Self-Supervised Vision Encoders AnimateDiff: Animate Your Personalized Text-to-Image Diffusion Models without Specific Tuning

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-04T18:58:10.908196Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:58:10.908196Z digest=sha256:eb2c2d8711bcf9854467d061292daa68deedca6bb34a71aab1a2cb4b4abf07d8

Observation f57d9569-41e4-4ca2-835d-ddd1a8fb5616 · outbound

This paper cites LTX-Video: Realtime Video Latent Diffusion.

Improving Video Diffusion Transformer Training by Multi-Feature Fusion and Alignment from Self-Supervised Vision Encoders LTX-Video: Realtime Video Latent Diffusion

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-04T18:58:10.913793Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:58:10.913793Z digest=sha256:7f6a7c2584c2f0a17dc9ce248c19fe1f296a27762ed88972565f3ab39fa2cbb9

Observation c6d61e40-6b87-49a5-8bfd-613e72c48ce7 · outbound

This paper cites JOG3R: Towards 3D-Consistent Video Generators.

Improving Video Diffusion Transformer Training by Multi-Feature Fusion and Alignment from Self-Supervised Vision Encoders JOG3R: Towards 3D-Consistent Video Generators

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-04T18:58:10.926506Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:58:10.926506Z digest=sha256:986c5bf4ec296238e29d4a55bd2d5076fd7d3ff265c5e041e1eeefd76faad5e7

Observation 7efd77b1-8cd2-4244-8b9d-ae20be33c484 · outbound

This paper cites Track4Gen: Teaching Video Diffusion Models to Track Points Improves Video Generation.

Improving Video Diffusion Transformer Training by Multi-Feature Fusion and Alignment from Self-Supervised Vision Encoders Track4Gen: Teaching Video Diffusion Models to Track Points Improves Video Generation

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-04T18:58:10.931053Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:58:10.931053Z digest=sha256:e3b741f6b9faad332cc6263c7eee7dbeddea29168d24194453e995ea8da5fa8a

Observation b4590f58-eee7-4510-981c-5546c821b0ba · outbound

This paper cites Exploring Temporally-Aware Features for Point Tracking.

Improving Video Diffusion Transformer Training by Multi-Feature Fusion and Alignment from Self-Supervised Vision Encoders Exploring Temporally-Aware Features for Point Tracking

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-04T18:58:10.935188Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:58:10.935188Z digest=sha256:92b0dcec2f408bcaad9878ff2d7d539ddf7be10c1f5d1589461a5618a970fc72

Observation a74e50d2-734c-49bc-bfd3-f7faa0312de4 · outbound

This paper cites HunyuanVideo: A Systematic Framework For Large Video Generative Models.

Improving Video Diffusion Transformer Training by Multi-Feature Fusion and Alignment from Self-Supervised Vision Encoders HunyuanVideo: A Systematic Framework For Large Video Generative Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-04T18:58:10.939891Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:58:10.939891Z digest=sha256:4cd1edbac1da98871aec6f8ac05c6b63d9aa230191c2d33da57dee582883cef9

Observation cb39bf73-02a7-4928-9c5d-776b76021199 · outbound

This paper cites Flow Straight and Fast: Learning to Generate and Transfer Data with Rectified Flow.

Improving Video Diffusion Transformer Training by Multi-Feature Fusion and Alignment from Self-Supervised Vision Encoders Flow Straight and Fast: Learning to Generate and Transfer Data with Rectified Flow

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-04T18:58:10.948812Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:58:10.948812Z digest=sha256:2cae23a9d2678218236a61080d1b86f65e3ec6249eca9f176279d04059176ee6

Observation 70974366-1796-44bb-9adf-73557c0f3467 · outbound

This paper cites Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model.

Improving Video Diffusion Transformer Training by Multi-Feature Fusion and Alignment from Self-Supervised Vision Encoders Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-04T18:58:10.953460Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:58:10.953460Z digest=sha256:1d3374cabc35dcbe6193ac206c71f530d6969d95fa5acfd6d46f8d689d5c495d

Observation 20fbdfa8-1737-4a96-a3c7-1e36bb9987ea · outbound

This paper cites Latte: Latent Diffusion Transformer for Video Generation.

Improving Video Diffusion Transformer Training by Multi-Feature Fusion and Alignment from Self-Supervised Vision Encoders Latte: Latent Diffusion Transformer for Video Generation

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-04T18:58:10.957915Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:58:10.957915Z digest=sha256:2500ee65718de71d9090e79666ba61653c273e5019b520965226741372ac7292

Observation dab65f3f-0119-4ca4-8306-636dee8c7695 · outbound

This paper cites OpenVid-1M: A Large-Scale High-Quality Dataset for Text-to-video Generation.

Improving Video Diffusion Transformer Training by Multi-Feature Fusion and Alignment from Self-Supervised Vision Encoders OpenVid-1M: A Large-Scale High-Quality Dataset for Text-to-video Generation

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-04T18:58:10.962343Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:58:10.962343Z digest=sha256:a01df38a5342439cab6cfe02e5b40dedd917a587dd150475dd818512e4687c98

Observation a84fc41f-d59e-4365-91d8-4c89d2d02fec · outbound

This paper cites SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis.

Improving Video Diffusion Transformer Training by Multi-Feature Fusion and Alignment from Self-Supervised Vision Encoders SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-04T18:58:10.967014Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:58:10.967014Z digest=sha256:4c1caab49ae4562dda43b676b418592238d66f2dd4d72307646b759d516e0bba

Observation ed61d7da-adb1-4e8f-bfea-3a36ff56f536 · outbound

This paper cites Movie Gen: A Cast of Media Foundation Models.

Improving Video Diffusion Transformer Training by Multi-Feature Fusion and Alignment from Self-Supervised Vision Encoders Movie Gen: A Cast of Media Foundation Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-04T18:58:10.972403Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:58:10.972403Z digest=sha256:b1ca2dcf157b14bad4c970e7373ccd9907d6c74bb4f5631db5b5519ffdec6a1e

Observation 0cd651a4-d122-43a5-9c68-3bcabeb37969 · outbound

This paper cites SAM 2: Segment Anything in Images and Videos.

Improving Video Diffusion Transformer Training by Multi-Feature Fusion and Alignment from Self-Supervised Vision Encoders SAM 2: Segment Anything in Images and Videos

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-04T18:58:10.976812Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:58:10.976812Z digest=sha256:5d462afbd42deb36512dec61d1af18a2c4174b82ea4534cc2d87a01ef7996ea4

Observation a14a2fc1-8be7-4e86-8bcf-a24956c4e526 · outbound

This paper cites FaceForensics: A Large-scale Video Dataset for Forgery Detection in Human Faces.

Improving Video Diffusion Transformer Training by Multi-Feature Fusion and Alignment from Self-Supervised Vision Encoders FaceForensics: A Large-scale Video Dataset for Forgery Detection in Human Faces

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-04T18:58:10.981630Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:58:10.981630Z digest=sha256:04f25b25115fafd7b053cf4f5353a1c72686dd28f5fe9f3ec0fb6c014af4ee6b

Observation e850c45d-85ec-453b-9861-d6de8c5f9c97 · outbound

This paper cites Make-A-Video: Text-to-Video Generation without Text-Video Data.

Improving Video Diffusion Transformer Training by Multi-Feature Fusion and Alignment from Self-Supervised Vision Encoders Make-A-Video: Text-to-Video Generation without Text-Video Data

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-04T18:58:10.986770Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:58:10.986770Z digest=sha256:f1cdbacb0d00d83d8b814f105689f0445ece11e04b791c0bde247bc9fb95d039

Observation 023655bb-12a9-472c-b6ef-6cfd1a033d53 · outbound

This paper cites arXiv:2502.06755.

Improving Video Diffusion Transformer Training by Multi-Feature Fusion and Alignment from Self-Supervised Vision Encoders arXiv:2502.06755

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-04T18:58:11.000548Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:58:11.000548Z digest=sha256:b006d6c5408311c5f0bf2a9d5fe5c55fa2a0c7ef905b9bc62b3f32327673577c

Observation 32143b16-82f4-4cd2-955b-d7be0b1c4bfa · outbound

This paper cites DUSt3R: Geometric 3D Vision Made Easy.

Improving Video Diffusion Transformer Training by Multi-Feature Fusion and Alignment from Self-Supervised Vision Encoders DUSt3R: Geometric 3D Vision Made Easy

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-04T18:58:11.013763Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:58:11.013763Z digest=sha256:5dd1e97d9c462fc1286f08f70cd002fbf07d00189b476df456e22b7606082e48

Observation 4474b6c3-8f6c-4003-add1-d6376784c4be · outbound

This paper cites InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation.

Improving Video Diffusion Transformer Training by Multi-Feature Fusion and Alignment from Self-Supervised Vision Encoders InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-04T18:58:11.018159Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:58:11.018159Z digest=sha256:e1c1936232eb30c6224b8eda617236b17fb4cbf41451db2b4b172c65552f23fe

Observation 0caf6e20-38e6-4699-8cc2-643cf12b1d78 · outbound

This paper cites CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer.

Improving Video Diffusion Transformer Training by Multi-Feature Fusion and Alignment from Self-Supervised Vision Encoders CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-04T18:58:11.022411Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:58:11.022411Z digest=sha256:5f2e4e4879c44c1f3e2973683e5c617dac5804c85f2f37fbbe1764ef39d8d6c7

Observation 0ee73c7b-dd24-4c7e-af27-0e254d57ea36 · outbound

This paper cites Representation Alignment for Generation: Training Diffusion Transformers Is Easier Than You Think.

Improving Video Diffusion Transformer Training by Multi-Feature Fusion and Alignment from Self-Supervised Vision Encoders Representation Alignment for Generation: Training Diffusion Transformers Is Easier Than You Think

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-04T18:58:11.026989Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:58:11.026989Z digest=sha256:6d477d9715066cc6577635ebdb89aab0f87e285f60aad19ce1ed35d829a1e2a1

Observation acaa8341-38dc-40b5-9112-76a1cd7fe508 · outbound

This paper cites MonST3R: A Simple Approach for Estimating Geometry in the Presence of Motion.

Improving Video Diffusion Transformer Training by Multi-Feature Fusion and Alignment from Self-Supervised Vision Encoders MonST3R: A Simple Approach for Estimating Geometry in the Presence of Motion

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-04T18:58:11.031336Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:58:11.031336Z digest=sha256:1ee0dda4ac25c41f0c2238bf75ab3dac035617d9b3180018aeec42e9bf0c8174

Observation 30b5a8eb-6250-443c-88e6-b9ba1b99d06c · outbound

This paper cites Allegro: Open the Black Box of Commercial-Level Video Generation Model.

Improving Video Diffusion Transformer Training by Multi-Feature Fusion and Alignment from Self-Supervised Vision Encoders Allegro: Open the Black Box of Commercial-Level Video Generation Model

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-04T18:58:11.035409Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:58:11.035409Z digest=sha256:a08575b8ab7257f09da5442cca4e08803e665e1a6cf61b2749a4a2a0f5017b1d

Observation e0e99ed2-34da-4416-a43f-9e67902b292c · outbound

This paper cites an unresolved cited work.

Improving Video Diffusion Transformer Training by Multi-Feature Fusion and Alignment from Self-Supervised Vision Encoders Unresolved cited work

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-04T18:58:11.040371Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:58:11.040371Z digest=sha256:5740a87c188b7eed380d7307994121fced6897a78b2ffd42f3c5b28bb1abc025

Observation d80ade39-1e31-4951-a917-832cf523219c · outbound

This paper cites We follow the same protocol as SkyTimeLapse, using the training split for both model training and metric evaluations.

Improving Video Diffusion Transformer Training by Multi-Feature Fusion and Alignment from Self-Supervised Vision Encoders We follow the same protocol as SkyTimeLapse, using the training split for both model training and metric evaluations

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-04T18:58:11.044648Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:58:11.044648Z digest=sha256:9b7f59bc679f9a4934f090a60628509301516062350b009c45ee34cb49202eee

Observation b42b5a19-1675-4d4e-b843-4317b198fca6 · outbound

This paper cites an unresolved cited work.

Improving Video Diffusion Transformer Training by Multi-Feature Fusion and Alignment from Self-Supervised Vision Encoders Unresolved cited work

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-04T18:58:11.049572Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:58:11.049572Z digest=sha256:53f1e52f88cbf955c5a23b81668b547e41be5b373a3a78160e02d7d1b435c519

Observation efe36236-3ec3-44db-89c3-4ac383d3aee4 · outbound

This paper cites UCF101: A Dataset of 101 Human Actions Classes From Videos in The Wild.

Improving Video Diffusion Transformer Training by Multi-Feature Fusion and Alignment from Self-Supervised Vision Encoders UCF101: A Dataset of 101 Human Actions Classes From Videos in The Wild

Reference 2012

Resolution
unresolved
no resolver link, observed 2026-08-04T18:58:10.996179Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:58:10.996179Z digest=sha256:9c8f8a1d14df91ebba9627af06abde17885db63a6a013c753a0e5e604f4392b1

Observation 4ad22f3d-e6d7-46ea-903b-8b16bd825241 · outbound

This paper cites Learning Spatiotemporal Features with 3D Convolutional Networks.

Improving Video Diffusion Transformer Training by Multi-Feature Fusion and Alignment from Self-Supervised Vision Encoders Learning Spatiotemporal Features with 3D Convolutional Networks

Reference 2015

Resolution
unresolved
no resolver link, observed 2026-08-04T18:58:11.004855Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:58:11.004855Z digest=sha256:3d81df201d35e61faf567f60eefed81e516d7070ea6b13f7cbcc4185ca9c4144

Observation 55714084-7ff2-4085-814b-b1580248107e · outbound

This paper cites GANs Trained by a Two Time-Scale Update Rule Converge to a Local Nash Equilibrium.

Improving Video Diffusion Transformer Training by Multi-Feature Fusion and Alignment from Self-Supervised Vision Encoders GANs Trained by a Two Time-Scale Update Rule Converge to a Local Nash Equilibrium

Reference 2018

Resolution
unresolved
no resolver link, observed 2026-08-04T18:58:10.922490Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:58:10.922490Z digest=sha256:8ab891905ea334e8359735e3cba084c100c6606e4e13b1de48f1e004f42170e0

Observation 57b70c0c-7f6c-4230-b65a-1b5ff02f817b · outbound

This paper cites Towards Accurate Generative Models of Video: A New Metric & Challenges.

Improving Video Diffusion Transformer Training by Multi-Feature Fusion and Alignment from Self-Supervised Vision Encoders Towards Accurate Generative Models of Video: A New Metric & Challenges

Reference 2019

Resolution
unresolved
no resolver link, observed 2026-08-04T18:58:11.009224Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:58:11.009224Z digest=sha256:12d9890a854a24773da864f30df30d92e1895f6f037490f08a7f698a53cb5adc

Observation 02cf6cb7-f7c8-46e1-8125-756c1754cef0 · outbound

This paper cites Denoising Diffusion Implicit Models.

Improving Video Diffusion Transformer Training by Multi-Feature Fusion and Alignment from Self-Supervised Vision Encoders Denoising Diffusion Implicit Models

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-04T18:58:10.991325Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:58:10.991325Z digest=sha256:5290839537dbf08c987844027cef1997adfa5a44bea8eba584a5f2907ab73abd

Observation 2987a97f-4eed-45ad-af7d-840e56491d7f · outbound

This paper cites Masked Autoencoders Are Scalable Vision Learners.

Improving Video Diffusion Transformer Training by Multi-Feature Fusion and Alignment from Self-Supervised Vision Encoders Masked Autoencoders Are Scalable Vision Learners

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-04T18:58:10.918242Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:58:10.918242Z digest=sha256:8f40bdcce4d49fb6bde40f1186c6d7761c51e64fea6a4d36ab787ef77cbffcb6

Observation 6f7e5855-ea2d-4866-abef-f2ca3b8b97a4 · outbound

This paper cites Flow Matching for Generative Modeling.

Improving Video Diffusion Transformer Training by Multi-Feature Fusion and Alignment from Self-Supervised Vision Encoders Flow Matching for Generative Modeling

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-04T18:58:10.944247Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:58:10.944247Z digest=sha256:7664b7c0f903a165a40a2baed2806354b6fd312f78d6d513b77d88f30ddb6531

Observation 26b0d4e0-4b54-4a25-8570-7882cd0b2820 · outbound

This paper cites PixArt-$\alpha$: Fast Training of Diffusion Transformer for Photorealistic Text-to-Image Synthesis.

Improving Video Diffusion Transformer Training by Multi-Feature Fusion and Alignment from Self-Supervised Vision Encoders PixArt-$\alpha$: Fast Training of Diffusion Transformer for Photorealistic Text-to-Image Synthesis

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-04T18:58:10.888699Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:58:10.888699Z digest=sha256:2bdeb2d037d634d4e67bc7124f155361b9b562c317dd564d604ef1b918e0eecc

Observation 9109bb9a-2641-4b02-b1c2-bdea3a36d576 · outbound

This paper cites Scaling Rectified Flow Transformers for High-Resolution Image Synthesis.

Improving Video Diffusion Transformer Training by Multi-Feature Fusion and Alignment from Self-Supervised Vision Encoders Scaling Rectified Flow Transformers for High-Resolution Image Synthesis

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-04T18:58:10.898380Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:58:10.898380Z digest=sha256:6c3e60191b31a1800c7ba41b6a48617732498442cd786275fede589ee64ff500

Observation 38f8dc10-0612-4b39-8cff-48529173ae1b · outbound

This paper cites VideoJAM: Joint Appearance-Motion Representations for Enhanced Motion Generation in Video Models.

Improving Video Diffusion Transformer Training by Multi-Feature Fusion and Alignment from Self-Supervised Vision Encoders VideoJAM: Joint Appearance-Motion Representations for Enhanced Motion Generation in Video Models

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-04T18:58:10.883059Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:58:10.883059Z digest=sha256:1d2b24a9aab81e8af3e46ee1e35c894005a7b98dd9b02681e727da8fd3d21dcd

Pith citing papers

No inbound Pith citation observations are available.