Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-19T08:02:23.002090Z
Paper Citation Record · LEDGER
As of 5 August 2026, this Paper Citation Record lists 100 of 296 outbound references and 56 inbound Pith citation observations for arXiv:2502.10248.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-19T08:02:23.002090Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-05T06:32:48.257954+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-05T17:45:17.186768Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-08-05T02:28:24.338817Z
100 of 296 outbound references displayed
External citation measurements
0
pith, observed 2026-08-05T02:28:24.338817Z
Observation e3db04c6-63d9-40ae-b163-e2319fa8a309 · outbound
Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model Video generation models as world simulators
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 1b51e1fb-457a-4895-8318-b1c0ca725923 · outbound
Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model Unresolved cited work
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 88567e7a-6b3e-4fd1-b6c9-75ce464fcd7b · outbound
Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model Unresolved cited work
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 57ba1ca1-fe51-44d4-a7a6-6b92721c1d77 · outbound
Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model Unresolved cited work
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation d0afb639-d642-455f-9e5a-ffb9b0a4d967 · outbound
Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model Gen-3 alpha
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation ce52b633-263e-4d43-b961-5582384821aa · outbound
Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model HunyuanVideo: A Systematic Framework For Large Video Generative Models
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 9726da56-2ba2-4d0d-b533-89429094b715 · outbound
Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model Open-sora: Democratizing efficient video production for all
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 6db8b2d6-a5b2-4c8e-8364-99b4949da35a · outbound
Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model Scaling Rectified Flow Transformers for High-Resolution Image Synthesis
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 1dcf930d-2b7a-4ff3-82b0-f0adb14fe80e · outbound
Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model Scalable Diffusion Models with Transformers
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 61cba769-5fc5-465b-898d-433d921fd919 · outbound
Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model Movie Gen: A Cast of Media Foundation Models
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 424b6607-d5d1-489b-9526-98c0a24c22ae · outbound
Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model Language model beats diffusion - tokenizer is key to visual generation
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 6f1d706e-5be6-48ce-a7eb-88a2acd01955 · outbound
Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model Cosmos World Foundation Model Platform for Physical AI
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 097c1c8e-d302-4d82-a71c-7411e8f83ff5 · outbound
Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model WF-VAE: Enhancing Video VAE by Wavelet-Driven Energy Flow for Latent Video Diffusion Model
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 2fdde92b-9479-49f6-a815-d659b52d27e6 · outbound
Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model Deep compression autoencoder for efficient high-resolution diffusion models
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 7e75ae10-ae45-4a05-8cbd-7c4ef9fdd5d0 · outbound
Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model Flow Matching for Generative Modeling
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 93a71d0d-a75f-49d1-ac8b-be50586c74c0 · outbound
Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model Hunyuan-DiT: A Powerful Multi-Resolution Diffusion Transformer with Fine-Grained Chinese Understanding
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 9f1b91b0-a607-4b5d-bf49-533c2766421b · outbound
Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model Train Short, Test Long: Attention with Linear Biases Enables Input Length Extrapolation
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation a122dd90-5703-4f45-8bdd-e3911295351b · outbound
Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model PixArt-$\alpha$: Fast Training of Diffusion Transformer for Photorealistic Text-to-Image Synthesis
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 51c3c749-932a-4e8f-b60a-eba91c382e48 · outbound
Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model RoFormer: Enhanced Transformer with Rotary Position Embedding
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 53b383e3-dd01-40b4-85ff-ae94aef16b22 · outbound
Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model Training language models to follow instructions with human feedback
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 4cac45e3-d861-4402-a285-31e63460f2d1 · outbound
Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model Deep reinforcement learning from human preferences
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation e8c9171a-e6dc-43de-aa6b-882ae4a670fe · outbound
Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model Direct preference optimization: Your language model is secretly a reward model
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 249d6c06-2337-4cda-b5f0-a620e157a6e2 · outbound
Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model Diffusion model alignment using direct preference optimization
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 390bbff3-0aca-45be-970a-fc41bf9c2ed6 · outbound
Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model Using human feedback to fine-tune diffusion models without any reward model
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 436ba729-9325-4305-8163-eb030321a751 · outbound
Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model Flow Straight and Fast: Learning to Generate and Transfer Data with Rectified Flow
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation cfd90426-d0ae-43e4-a433-3205fda778ad · outbound
Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model Improving the Training of Rectified Flows
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 374ee2e8-dbe3-4d97-8196-241ec31dd8c1 · outbound
Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model Efficient large-scale language model training on gpu clusters using megatron-lm
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 3fb1ec26-f61a-4718-a126-0fbc3e316c31 · outbound
Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model Reducing activation recomputation in large transformer models
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 6e5151fa-990e-414d-a290-97ac6844e0d4 · outbound
Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model Ring Attention with Blockwise Transformers for Near-Infinite Context
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 78b559a4-8f39-4cc0-a01f-3c47b76d14b5 · outbound
Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model DeepSpeed Ulysses: System Optimizations for Enabling Training of Extreme Long Sequence Transformer Models
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 59c8ac4b-0e32-4482-b50f-29a59eaf51e4 · outbound
Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model Zero: Memory optimizations toward training trillion parameter models
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 92699452-73e7-45cd-947a-6e2afc3ac75f · outbound
Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model Disttrain: Addressing model and data heterogeneity with disaggregated training for multimodal large language models
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 5afcdcbc-5d7e-4568-a5f8-eedf8c1efd40 · outbound
Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model Channels Last Memory Format in PyTorch
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 96ebe783-8ae8-4e0a-9c07-6b88701ac625 · outbound
Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model Jordan, and Ion Stoica
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 3b3631c5-b79b-4018-89aa-00fbec5f72e2 · outbound
Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model Pytorch rpc: Distributed deep learning built on tensor-optimized remote procedure calls
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 1597fe2b-4566-4862-95e8-7fe26d69ce6c · outbound
Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model Mooncake: A KVCache-centric Disaggregated Architecture for LLM Serving
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 33fa9489-9a01-4f62-aee0-54cce38f314f · outbound
Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model The llama 3 herd of models, April 2024
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 6220c3a1-967c-444b-89f2-dd4e2e8bad5c · outbound
Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model PySceneDetect
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 4a718b9d-7051-47c9-abec-45dcd19d65de · outbound
Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model Unresolved cited work
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 698c0a06-2230-4b30-a24f-5c333ae0832a · outbound
Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model Panda-70m: Captioning 70m videos with multiple cross-modality teachers
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation b69d111d-f1b2-40b2-9ff3-2723e9c50d31 · outbound
Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model Laion-5b: An open large-scale dataset for training next generation image-text models
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation f3ed22ab-8066-4c94-948f-362cbeeefc68 · outbound
Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model Clip-based nsfw detector
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 7de259bf-3738-4e5d-b8c2-a5445dc425c0 · outbound
Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model Efficientnet: Rethinking model scaling for convolutional neural networks
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation a364f9ec-d3d8-4035-b293-6d5c6951667c · outbound
Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model Paddleocr
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 7daad267-6ed3-44a5-b876-b5cda10d7c8d · outbound
Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model Unresolved cited work
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 81735cd3-50dd-4f7c-aeb3-0f7bfcfd589a · outbound
Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model Diatom autofocusing in brightfield microscopy: a comparative study
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 4ff1bcaa-a0ab-43ad-8a6c-765d98a938c5 · outbound
Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model Improving image generation with better captions
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 8ebdd587-9a73-4601-9a99-2bb5b964648d · outbound
Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model Some methods for classification and analysis of multivariate observations
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation f61c2c9b-4f03-49ce-a065-2a4b73c8d210 · outbound
Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model DSV: Exploiting Dynamic Sparsity to Accelerate Large-Scale Video DiT Training
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 34702d07-6542-437e-a503-2fe9521d526b · outbound
Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model Diffusion Forcing: Next-token Prediction Meets Full-Sequence Diffusion
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation afda7481-5ba0-4838-b530-1711a28a1ee6 · outbound
Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model LTX-Video: Realtime Video Latent Diffusion
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 345c7c81-7ca0-4f60-9dcd-030355ebec85 · outbound
Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model Taming Teacher Forcing for Masked Autoregressive Video Generation
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation f6b0c620-dc27-4df5-b667-8f25f95a4d3e · outbound
Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 383236cd-3173-4e19-a441-9fb768e7a307 · outbound
Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model 2024 , journal =
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 6bb8a284-d037-4930-b7bc-82170cd6f73a · outbound
Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model 2024 , eprint=
Reference 59
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 1e4c5300-0089-4db7-aaf8-9dda2e70d401 · outbound
Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model 2024 , eprint=
Reference 60
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation b2d931d3-ba4b-4117-8aab-26e23e1a1725 · outbound
Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model 2025 , eprint=
Reference 61
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 56820fb5-a3f8-4e25-8b24-51a84497a81d · outbound
Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model Computer Science
Reference 62
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 04445759-84d1-4007-a6b6-aa4f445b0e57 · outbound
Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model 2024 , eprint=
Reference 63
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 9e582ab3-6d34-4dbf-b892-555e7ecb7c85 · outbound
Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model 2023 , eprint=
Reference 64
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation ec80df0b-057d-40ac-bd55-801dabc53a30 · outbound
Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model 2024 , eprint=
Reference 65
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 5ee362d6-1eff-4af4-a14c-8976d9597019 · outbound
Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model 2024 , eprint=
Reference 66
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 938b2f6f-f49b-42f9-a20e-594e48df9a9c · outbound
Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer
Reference 67
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 208833bf-b0ca-4eef-a99c-11d0bc2de12a · outbound
Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model 2024 , url =
Reference 68
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 7475e662-ce52-48f3-97e8-f00428f6dfd3 · outbound
Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model Open-Sora Plan: Open-Source Large Video Generation Model
Reference 69
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation c90974cd-bf48-4d76-9ec8-ca9582573bcc · outbound
Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model 2025 , eprint=
Reference 70
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 9aea7952-175a-49e6-bf36-687b94cbb0c9 · outbound
Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model 2024 , journal =
Reference 71
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation bde6ba1a-2c12-43db-a679-a77e55f9fb45 · outbound
Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model 2024 , journal =
Reference 72
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 9e817465-1af0-458a-8e07-2ac2c852c9de · outbound
Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model 2024 , journal =
Reference 73
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation c86d4304-38f4-4c2a-8e98-6c088b0c8ca9 · outbound
Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model 2024 , journal =
Reference 74
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation ed2c5a0d-8d5a-4d85-948a-e708b64277b1 · outbound
Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model 2025 , eprint=
Reference 75
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 5cfb6c52-bd07-458c-970a-895d3de02a3e · outbound
Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model 2023 , eprint=
Reference 76
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 582de9a2-1b1b-41c2-a53b-f398898b442e · outbound
Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model and Ilharco, Gabriel and Song, Shuran and Kollar, Thomas and Carmon, Yair and Dave, Achal and Heckel, Reinhard and Muennighoff, Niklas and Schmidt, Ludwig , title=
Reference 77
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 5169f380-c5e0-4def-bdd1-48a43e8457ca · outbound
Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model 2023 , eprint=
Reference 78
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 05f79cd2-3e60-472a-a480-73b736354cae · outbound
Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model PaLM 2 Technical Report
Reference 79
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation c6ceac81-6d16-4aeb-98e0-477ea8bbe813 · outbound
Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model CCNet: Extracting High Quality Monolingual Datasets from Web Crawl Data
Reference 80
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation f66bfdee-d959-4ca7-9f0a-fdb399b09517 · outbound
Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model LLaMA: Open and Efficient Foundation Language Models
Reference 81
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 9a4bdb57-52db-44ba-84ea-5b98a6f043b1 · outbound
Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model Mistral 7B
Reference 82
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation cf5eedea-4314-42f7-bb11-311f44ac5782 · outbound
Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model The Twelfth International Conference on Learning Representations , year=
Reference 83
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 42719aa2-0294-4d5b-8abf-a17cc92315af · outbound
Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model Unresolved cited work
Reference 84
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 52617cc4-a03e-4774-afd0-4e755f00c81d · outbound
Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model StarCoder: may the source be with you! , journal =
Reference 85
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 80514e52-fa1d-47df-a965-6393ddb5fc42 · outbound
Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model 2024 , eprint=
Reference 86
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 2e8c8a6a-7384-4052-9e62-380090c3f296 · outbound
Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model Gemma: Open Models Based on Gemini Research and Technology
Reference 87
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 20f033f6-595f-4f2b-9f65-8e88f1537a10 · outbound
Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model Qwen Technical Report
Reference 88
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation f8a02304-ef88-4a23-a199-55336e0b03dd · outbound
Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model Unresolved cited work
Reference 89
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation c1a9ab32-f2c9-404c-a30e-e0f51168c90c · outbound
Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model Microsoft Research Blog , year=
Reference 90
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 375cea56-7ee4-46c8-8e7d-f8386e8f1a49 · outbound
Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model DeepSeek LLM: Scaling Open-Source Language Models with Longtermism
Reference 91
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 2c5cd385-a1f4-4238-929f-7be6eae0d30b · outbound
Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model Code Llama: Open Foundation Models for Code
Reference 92
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation d37b6d8b-f2e7-4ca1-8e66-2b13acfb6b95 · outbound
Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model Advances in Neural Information Processing Systems , volume=
Reference 93
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 4f79123f-43bb-42fb-a6e8-7b6ada5e23df · outbound
Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model Llemma: An Open Language Model For Mathematics
Reference 94
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 3179ab27-421d-4ecb-b699-ce7525ad7db8 · outbound
Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model InternLM-Math: Open Math Large Language Models Toward Verifiable Reasoning
Reference 95
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation d4b983ec-f9d8-4e09-9f41-7173e8d0315d · outbound
Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model Let's Verify Step by Step
Reference 96
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation ed18654e-3653-4deb-b6da-9f7ac6f729de · outbound
Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model The Twelfth International Conference on Learning Representations , year=
Reference 97
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 9f330a17-c10e-487b-8698-8375c9dbe106 · outbound
Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model Large-scale Dataset Pruning with Dynamic Uncertainty
Reference 98
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation d345aa5e-0f29-4cda-83f2-4a05c6132b5c · outbound
Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model NIPS , volume=
Reference 99
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 755a65f1-c9ee-420e-b1de-c12ed84ad7d4 · outbound
Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model From Quantity to Quality: Boosting LLM Performance with Self-Guided Data Selection for Instruction Tuning
Reference 100
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation e19c190a-a269-4b39-bc60-e79be703cff5 · outbound
Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model One-Shot Learning as Instruction Data Prospector for Large Language Models
Reference 101
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 0b3012be-c414-4d6c-81bb-f3c4215a930a · outbound
Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model Self-Evolved Diverse Data Sampling for Efficient Instruction Tuning
Reference 102
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 510cf5ae-db2d-4c7d-a76b-282746e4a2eb · outbound
Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model ICLR , year=
Reference 103
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 85d0480e-642e-47a8-86f7-1f6c1e09f8b1 · outbound
Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model ICLR , year=
Reference 104
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 44c1afa1-920b-4ea5-98cf-fd32047b824f · inbound
Wan: Open and Advanced Large-Scale Video Generative Models Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 31dcae6c-86f1-42e0-880c-09d71cdcfaf0 · inbound
VBench-2.0: Advancing Video Generation Benchmark Suite for Intrinsic Faithfulness Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model
Reference 62
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation eed6a369-57ed-43b2-a4a0-53823b0bb2fe · inbound
MAGI-1: Autoregressive Video Generation at Scale Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 0d4a097e-ebf6-431b-9307-90e7333825a0 · inbound
GenHSI: Controllable Generation of Human-Scene Interaction Videos Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model
Reference 61
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 6b1207d8-4f4c-4907-b48b-14a385e92439 · inbound
Listener-Rewarded Thinking in VLMs for Image Preferences Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 8802044d-ba60-45f8-967f-f8ae9eaab0df · inbound
Waver: Wave Your Way to Lifelike Video Generation Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 566e6d69-9a7f-4b12-a148-f183e1019dc6 · inbound
UniVerse-1: Unified Audio-Video Generation via Stitching of Experts Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e5445f7e-d9e2-48ea-89bf-662cbd1b64c2 · inbound
RewardDance: Reward Scaling in Visual Generation Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 70974366-1796-44bb-9adf-73557c0f3467 · inbound
Improving Video Diffusion Transformer Training by Multi-Feature Fusion and Alignment from Self-Supervised Vision Encoders Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b966e68f-0b2e-4e8e-af02-c24edfaacfdd · inbound
Enhancing Physical Plausibility in Video Generation by Reasoning the Implausibility Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 8a4bf5bb-9ddc-4ecb-9c1e-0b0b518987e1 · inbound
UniVideo: Unified Understanding, Generation, and Editing for Videos Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e52145d8-551f-44ff-963f-decce4ea2148 · inbound
HunyuanVideo 1.5 Technical Report Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 78382ff9-3617-408b-b8a7-f2b23c649fe4 · inbound
Delving into Latent Spectral Biasing of Video VAEs for Superior Diffusability Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 58b1a2bb-31ae-4566-bf77-29274feb2d93 · inbound
VideoASMR-Bench: Can AI-Generated ASMR Videos Fool VLMs and Humans? Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 98638ac3-139f-431d-9cf9-f47b4921b9ca · inbound
Self-transcendence: Is External Feature Guidance Indispensable for Accelerating Diffusion Transformer Training? Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6c7e2aae-f7a7-4913-91ca-698675b9abf6 · inbound
Beyond Rigid: Benchmarking Non-Rigid Video Editing Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model
Reference 379
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9edbeec9-48e9-47c3-9ad7-030279c654b5 · inbound
CG-MLLM: Captioning and Generating 3D content via Multi-modal Large Language Models Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 9deb6519-0c9e-448f-bdd1-1f7fa87fdf40 · inbound
SynthForensics: Benchmarking and Evaluating People-Centric Synthetic Video Deepfakes Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 9c7ac717-3e3a-48fb-9ebb-72ff1e99671f · inbound
EchoTorrent: Towards Swift, Sustained, and Streaming Multi-Modal Video Generation Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model
Reference 60
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 5112e65a-e406-4eb6-865b-585573014806 · inbound
Event-Driven Video Generation Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 11af68b9-8d46-46fb-895f-0186e44043c4 · inbound
ActionParty: Multi-Subject Action Binding in Generative Video Games Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 45280eac-eeb0-46f0-856c-4213cafa5701 · inbound
Evolution of Video Generative Foundations Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model
Reference 82
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation a0c1f447-6ba6-4461-a89f-a9e5ca911a87 · inbound
Efficient Video Diffusion Models: Advancements and Challenges Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model
Reference 100
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 6616cf1e-aa88-4c11-b187-2035ae3fb2cb · inbound
Motif-Video 2B: Technical Report Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 5e529286-57be-4923-bc51-63bd9064b1ce · inbound
Motif-Video 2B: Technical Report Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation ce6aff6d-cd09-4797-8266-a6ec2589d1e2 · inbound
DreamShot: Personalized Storyboard Synthesis with Video Diffusion Prior Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation f55539d9-13c8-43c1-9ab3-87e8da5cf939 · inbound
Leveraging Verifier-Based Reinforcement Learning in Image Editing Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 87934217-edb9-465f-930d-db6142090418 · inbound
Leveraging Verifier-Based Reinforcement Learning in Image Editing Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 0ce52856-1c69-46a6-804b-832e1d875b22 · inbound
Offline Preference Optimization for Rectified Flow with Noise-Tracked Pairs Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 84501807-5c5d-40ce-ab05-21253e441eaf · inbound
Qwen-Image-2.0 Technical Report Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 01c147af-d55c-4188-99ee-759ec01aca87 · inbound
HorizonDrive: Self-Corrective Autoregressive World Model for Long-horizon Driving Simulation Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation d427552a-f2fe-4715-a1fb-3be5be616e75 · inbound
HorizonDrive: Self-Corrective Autoregressive World Model for Long-horizon Driving Simulation Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 9a1d3f35-bb45-4a63-9302-acba56d47bde · inbound
Qwen-Image-VAE-2.0 Technical Report Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 658e9d79-8cae-437e-943c-f6038e0cc3e8 · inbound
HASTE: Training-Free Video Diffusion Acceleration via Head-Wise Adaptive Sparse Attention Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation b95885bf-e0f8-4cda-a132-16357f26e81a · inbound
MechVerse: Evaluating Physical Motion Consistency in Video Generation Models Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation bca9dbe6-be37-4f29-ad2b-7996904c7f08 · inbound
RAVEN: Real-time Autoregressive Video Extrapolation with Consistency-model GRPO Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation ecead28d-b9d4-4590-ab41-ecf0a83832a6 · inbound
AtlasVid: Efficient Ultra-High-Resolution Long Video Generation via Decoupled Global-Local Modeling Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 58ef28cc-f69a-4fb4-82e9-fea7321a3ea3 · inbound
Lance: Unified Multimodal Modeling by Multi-Task Synergy Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model
Reference 79
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 8e42cb5d-e659-4374-9e79-958876ff0e86 · inbound
Lance: Unified Multimodal Modeling by Multi-Task Synergy Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model
Reference 80
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 9f8f8f6d-c0f6-41ad-a266-038ebbde8907 · inbound
Bernini: Latent Semantic Planning for Video Diffusion Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 52033316-44b7-4c17-9bb6-4103ef7e26a0 · inbound
Future Forcing: Future-aware Training-free KV Cache Policy for Autoregressive Video Generation Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 2003d70b-23fc-4d66-9163-b9d369efd0a3 · inbound
Veda: Scalable Video Diffusion via Distilled Sparse Attention Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation bfa28da5-434f-435f-acdd-876ba781297a · inbound
Diffusing in the Right Space: A Systematic Study of Latent Diffusability Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 5b87e6b5-7b92-4ffb-b7bc-59df73892939 · inbound
ARGUS: Stacked Multi-View Identity Mosaic Injection for Subject-Preserving Video Generation Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation be1cef6d-5305-4b0b-8ec0-2c394241452d · inbound
PhysRAG: Enhancing Physics-Awareness in Video Generation via Retrieval-Augmented Generation Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 49a440ee-1c15-421f-aa6d-ff5c44ac9f06 · inbound
Bridging Video Understanding and Generation in a Unified Framework Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 48e00456-ad5a-47f8-9b9c-604d58d01147 · inbound
Ink3D: Sculpting 3D Assets with Extremely Complex Textures via Video Generative Models Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 87c8e050-6a90-45cb-9540-c7003dc306b4 · inbound
Arachne: Orchestrating Cascades for Efficient Text-to-Video Model Training Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation a949b09d-e383-4613-8c74-1d71f7095a61 · inbound
AlayaWorld: Long-Horizon and Playable Video World Generation Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 3b0ab594-a380-4434-8d80-9a12d2a319ca · inbound
Kaleido: Algorithm-Hardware Co-Design for Video Diffusion Transformers by Exploiting Latent Space Correlations Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4199c132-968a-48cf-9678-9358b7234195 · inbound
DiTango: Cost-Effective Parallel Diffusion Generation with Selective Attention State Reuse Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 377abb5b-dc37-4e3c-acf3-546a995b061c · inbound
Multi-Dimensional Quality Assessment for AI-Generated Human-Centric Videos: Dataset and Model Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation da3bfc9d-02b5-4098-9c30-bb6e38287aa2 · inbound
HarmoHOI: Harmonizing Appearance and 3D Motion for Multi-view Hand-Object Interaction Synthesis Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model
Reference 124
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3f4ba15e-323c-462c-bc24-32f845b45eeb · inbound
ABot-World-0: Infinite Interactive World Rollout on a Single Desktop GPU Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fb4d8cc3-cee3-48df-a927-b53def6cf275 · inbound
VideoCoCo: Code-as-CoT for Physically-Consistent Video Generation via an Agentic Dual-Engine System Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d08a136b-85b4-4c8a-a2c8-25ec90a6b775 · inbound
Retrieval-Driven Training-Free AI-Generated Video Attribution Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.