Pith. sign in

Paper Citation Record · LEDGER

Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model

As of 5 August 2026, this Paper Citation Record lists 100 of 296 outbound references and 56 inbound Pith citation observations for arXiv:2502.10248.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2502.10248 v3

Coverage vector

measured 100 of 296 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-19T08:02:23.002090Z

measured 156 of 156 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-05T06:32:48.257954+00:00

measured 56 of 56 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-05T17:45:17.186768Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-05T02:28:24.338817Z

Reference resolution

100 of 296 outbound references displayed

  • verified exact22
  • verified fuzzy54
  • unresolved7
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch17

External citation measurements

0
pith, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation e3db04c6-63d9-40ae-b163-e2319fa8a309 · outbound

This paper cites Video generation models as world simulators.

Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model Video generation models as world simulators

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T08:02:24.416179Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-19T08:02:23.002090Z digest=sha256:43414a5a631a32139caa0c8c803d68bfd3427f90392ebe503ce56f1eb1c7c78f

Observation 1b51e1fb-457a-4895-8318-b1c0ca725923 · outbound

This paper cites an unresolved cited work.

Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model Unresolved cited work

Reference 2

Resolution
unresolved
raw_fallback, observed 2026-05-19T08:02:24.425629Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-19T08:02:23.002090Z digest=sha256:08e62deec2f1077d6344cfd6526c5ab3faf3cd70ed124c38073090ed60a7cfdf

Observation 88567e7a-6b3e-4fd1-b6c9-75ce464fcd7b · outbound

This paper cites an unresolved cited work.

Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model Unresolved cited work

Reference 3

Resolution
unresolved
raw_fallback, observed 2026-05-19T08:02:24.430257Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-19T08:02:23.002090Z digest=sha256:49ba4aa031b3360ddc9b9af4434d5bcc853991eee1b7f70a578ccce2e27189a1

Observation 57ba1ca1-fe51-44d4-a7a6-6b92721c1d77 · outbound

This paper cites an unresolved cited work.

Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model Unresolved cited work

Reference 4

Resolution
unresolved
raw_fallback, observed 2026-05-19T08:02:24.440335Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-19T08:02:23.002090Z digest=sha256:4c154ee16621960c632acef063e254c643ead426d306945152998cffcedadc7a

Observation d0afb639-d642-455f-9e5a-ffb9b0a4d967 · outbound

This paper cites Gen-3 alpha.

Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model Gen-3 alpha

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T08:02:24.444978Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-19T08:02:23.002090Z digest=sha256:fcfd6b6f142747c7ed763b295b66e2cd21a81b299d01570d6ccf54b9226f418e

Observation ce52b633-263e-4d43-b961-5582384821aa · outbound

This paper cites HunyuanVideo: A Systematic Framework For Large Video Generative Models.

Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model HunyuanVideo: A Systematic Framework For Large Video Generative Models

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-05-19T08:02:23.931507Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-19T08:02:23.002090Z digest=sha256:73bf98753fdea98868682c5a81132a6be272f5463096d07cd0bde6726b3f8c5d

Observation 9726da56-2ba2-4d0d-b533-89429094b715 · outbound

This paper cites Open-sora: Democratizing efficient video production for all.

Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model Open-sora: Democratizing efficient video production for all

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T08:02:24.449409Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-19T08:02:23.002090Z digest=sha256:91c92721cb8a3404e2b199031f1395f9c74c0e9dea5d86badd0ea543a8d8fe77

Observation 6db8b2d6-a5b2-4c8e-8364-99b4949da35a · outbound

This paper cites Scaling Rectified Flow Transformers for High-Resolution Image Synthesis.

Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model Scaling Rectified Flow Transformers for High-Resolution Image Synthesis

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-05-19T08:02:24.057710Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-19T08:02:23.002090Z digest=sha256:62309cefa6f984b40b38bb85992a0f265269dc70db5e1f2b97a759a284e36ad1

Observation 1dcf930d-2b7a-4ff3-82b0-f0adb14fe80e · outbound

This paper cites Scalable Diffusion Models with Transformers.

Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model Scalable Diffusion Models with Transformers

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-05-19T08:02:23.796342Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-19T08:02:23.002090Z digest=sha256:40c83d1e2d93239330a5c4cfa43163935455aa5deef909352445e358995dfe06

Observation 61cba769-5fc5-465b-898d-433d921fd919 · outbound

This paper cites Movie Gen: A Cast of Media Foundation Models.

Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model Movie Gen: A Cast of Media Foundation Models

Reference 12

Resolution
metadata mismatch
local_arxiv, observed 2026-05-19T08:02:23.924392Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-19T08:02:23.002090Z digest=sha256:23155357aad09a4be2b29f421c79cdc610a6daa9d27c29464a9c37f6879a9b89

Observation 424b6607-d5d1-489b-9526-98c0a24c22ae · outbound

This paper cites Language model beats diffusion - tokenizer is key to visual generation.

Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model Language model beats diffusion - tokenizer is key to visual generation

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T08:02:24.453959Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-19T08:02:23.002090Z digest=sha256:17d1a2b94daf8666b61029768f472ba8fdba77277a7e8452e4e2a3f7b8bde2ab

Observation 6f1d706e-5be6-48ce-a7eb-88a2acd01955 · outbound

This paper cites Cosmos World Foundation Model Platform for Physical AI.

Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model Cosmos World Foundation Model Platform for Physical AI

Reference 14

Resolution
verified exact
local_arxiv, observed 2026-05-19T08:02:23.962171Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-19T08:02:23.002090Z digest=sha256:c874fe5913a35a21b9763fc7e0fbab41f0321dc3a07f7039444506516134f533

Observation 097c1c8e-d302-4d82-a71c-7411e8f83ff5 · outbound

This paper cites WF-VAE: Enhancing Video VAE by Wavelet-Driven Energy Flow for Latent Video Diffusion Model.

Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model WF-VAE: Enhancing Video VAE by Wavelet-Driven Energy Flow for Latent Video Diffusion Model

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-19T08:02:23.987814Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-19T08:02:23.002090Z digest=sha256:a9be3b8a081241ce1b45ba2a20a7977f7a555044587b9057fe902214509fdf8f

Observation 2fdde92b-9479-49f6-a815-d659b52d27e6 · outbound

This paper cites Deep compression autoencoder for efficient high-resolution diffusion models.

Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model Deep compression autoencoder for efficient high-resolution diffusion models

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T08:02:24.459607Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-19T08:02:23.002090Z digest=sha256:1542564cb9701d9ef40881059c68d89a9b2f04e225589b8f40a490b989b1bf15

Observation 7e75ae10-ae45-4a05-8cbd-7c4ef9fdd5d0 · outbound

This paper cites Flow Matching for Generative Modeling.

Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model Flow Matching for Generative Modeling

Reference 17

Resolution
verified exact
local_arxiv, observed 2026-05-19T08:02:24.071750Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-19T08:02:23.002090Z digest=sha256:d0e1bbbdb6dffe66ada2dd223a133672f3287f012dedfabb19cf7d5e256d62fb

Observation 93a71d0d-a75f-49d1-ac8b-be50586c74c0 · outbound

This paper cites Hunyuan-DiT: A Powerful Multi-Resolution Diffusion Transformer with Fine-Grained Chinese Understanding.

Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model Hunyuan-DiT: A Powerful Multi-Resolution Diffusion Transformer with Fine-Grained Chinese Understanding

Reference 18

Resolution
verified exact
local_arxiv, observed 2026-05-19T08:02:23.310121Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-19T08:02:23.002090Z digest=sha256:c6d0a6e9c8c2276cd105f1be557ce4a7e4a6c7dc24f94a6be1149fee11793db3

Observation 9f1b91b0-a607-4b5d-bf49-533c2766421b · outbound

This paper cites Train Short, Test Long: Attention with Linear Biases Enables Input Length Extrapolation.

Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model Train Short, Test Long: Attention with Linear Biases Enables Input Length Extrapolation

Reference 19

Resolution
metadata mismatch
local_arxiv, observed 2026-05-19T08:02:23.584841Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-19T08:02:23.002090Z digest=sha256:8f92c6448695ac76d3418e7e054f99b7b08038a2498cde2f7589a10566cc1f23

Observation a122dd90-5703-4f45-8bdd-e3911295351b · outbound

This paper cites PixArt-$\alpha$: Fast Training of Diffusion Transformer for Photorealistic Text-to-Image Synthesis.

Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model PixArt-$\alpha$: Fast Training of Diffusion Transformer for Photorealistic Text-to-Image Synthesis

Reference 20

Resolution
verified exact
local_arxiv, observed 2026-05-19T08:02:23.592279Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-19T08:02:23.002090Z digest=sha256:ecc55a1208c180c4581e669879f51722a2cf522711bda168ebb7bfb9750b7554

Observation 51c3c749-932a-4e8f-b60a-eba91c382e48 · outbound

This paper cites RoFormer: Enhanced Transformer with Rotary Position Embedding.

Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model RoFormer: Enhanced Transformer with Rotary Position Embedding

Reference 21

Resolution
verified exact
local_arxiv, observed 2026-05-19T08:02:23.671701Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-19T08:02:23.002090Z digest=sha256:9e18f9976fd2213d6ada0b6eb6d6ccc6c0709073246edb51b7e4dced87ab608f

Observation 53b383e3-dd01-40b4-85ff-ae94aef16b22 · outbound

This paper cites Training language models to follow instructions with human feedback.

Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model Training language models to follow instructions with human feedback

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T08:02:24.465394Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-19T08:02:23.002090Z digest=sha256:9374f186d965120a2c0b50290a7034e0071521cbf25678d084598bcae25f9ff5

Observation 4cac45e3-d861-4402-a285-31e63460f2d1 · outbound

This paper cites Deep reinforcement learning from human preferences.

Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model Deep reinforcement learning from human preferences

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T08:02:24.470401Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-19T08:02:23.002090Z digest=sha256:067c05fe13ba1edd35e1b633115116f07a55b06690542283bf9ae2c2eb49250f

Observation e8c9171a-e6dc-43de-aa6b-882ae4a670fe · outbound

This paper cites Direct preference optimization: Your language model is secretly a reward model.

Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model Direct preference optimization: Your language model is secretly a reward model

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T08:02:24.477172Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-19T08:02:23.002090Z digest=sha256:3592993355cebc6f5ef6392074402507a22a6ab72d63e4506ebabebd38821d95

Observation 249d6c06-2337-4cda-b5f0-a620e157a6e2 · outbound

This paper cites Diffusion model alignment using direct preference optimization.

Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model Diffusion model alignment using direct preference optimization

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T08:02:24.481525Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-19T08:02:23.002090Z digest=sha256:0787fc941f63099874fe51f933e9155a9836a71b8d1817f8b0c140573811b030

Observation 390bbff3-0aca-45be-970a-fc41bf9c2ed6 · outbound

This paper cites Using human feedback to fine-tune diffusion models without any reward model.

Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model Using human feedback to fine-tune diffusion models without any reward model

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T08:02:24.487178Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-19T08:02:23.002090Z digest=sha256:01a6e01f1afc2b0fe0ef9b2eaba2bf2a01140a2a18acd10fbc583dc8003c00f9

Observation 436ba729-9325-4305-8163-eb030321a751 · outbound

This paper cites Flow Straight and Fast: Learning to Generate and Transfer Data with Rectified Flow.

Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model Flow Straight and Fast: Learning to Generate and Transfer Data with Rectified Flow

Reference 27

Resolution
verified exact
local_arxiv, observed 2026-05-19T08:02:24.018974Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-19T08:02:23.002090Z digest=sha256:665bb48e9a57057e6088783c2d25400eba6ba01238e1781c5ec0f7fb8a5db99a

Observation cfd90426-d0ae-43e4-a433-3205fda778ad · outbound

This paper cites Improving the Training of Rectified Flows.

Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model Improving the Training of Rectified Flows

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-05-19T08:02:24.026542Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-19T08:02:23.002090Z digest=sha256:f0c4bbecee55355d235635d8c44a6caca470ab335bff80829f8a86540d060d49

Observation 374ee2e8-dbe3-4d97-8196-241ec31dd8c1 · outbound

This paper cites Efficient large-scale language model training on gpu clusters using megatron-lm.

Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model Efficient large-scale language model training on gpu clusters using megatron-lm

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T08:02:24.491468Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-19T08:02:23.002090Z digest=sha256:865637b59da841671b9404848433f137b495185d0ef5f08ae45064c660cb08cc

Observation 3fb1ec26-f61a-4718-a126-0fbc3e316c31 · outbound

This paper cites Reducing activation recomputation in large transformer models.

Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model Reducing activation recomputation in large transformer models

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T08:02:24.498241Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-19T08:02:23.002090Z digest=sha256:10163922a415a4b818b797a4542a4c7a89e1cc687d45fa5c5dc33414b8aaf317

Observation 6e5151fa-990e-414d-a290-97ac6844e0d4 · outbound

This paper cites Ring Attention with Blockwise Transformers for Near-Infinite Context.

Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model Ring Attention with Blockwise Transformers for Near-Infinite Context

Reference 31

Resolution
verified exact
local_arxiv, observed 2026-05-19T08:02:24.077796Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-19T08:02:23.002090Z digest=sha256:5e21e56cce505097c8b5798bef932c214bf811a81c425919617d49ef374726ca

Observation 78b559a4-8f39-4cc0-a01f-3c47b76d14b5 · outbound

This paper cites DeepSpeed Ulysses: System Optimizations for Enabling Training of Extreme Long Sequence Transformer Models.

Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model DeepSpeed Ulysses: System Optimizations for Enabling Training of Extreme Long Sequence Transformer Models

Reference 32

Resolution
verified exact
local_arxiv, observed 2026-05-19T08:02:23.291569Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-19T08:02:23.002090Z digest=sha256:239297c758bac4191b20012e4d7902649f9d0c6db4a79a217e0429fca14a95ff

Observation 59c8ac4b-0e32-4482-b50f-29a59eaf51e4 · outbound

This paper cites Zero: Memory optimizations toward training trillion parameter models.

Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model Zero: Memory optimizations toward training trillion parameter models

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T08:02:24.503153Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-19T08:02:23.002090Z digest=sha256:ace1dc4b631a8622170e7f845272d77c2acc2e01297c0fdc7585fb37c62bedc2

Observation 92699452-73e7-45cd-947a-6e2afc3ac75f · outbound

This paper cites Disttrain: Addressing model and data heterogeneity with disaggregated training for multimodal large language models.

Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model Disttrain: Addressing model and data heterogeneity with disaggregated training for multimodal large language models

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-05-19T08:02:23.564702Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-19T08:02:23.002090Z digest=sha256:26f1be4f1c0fc0bb8086a8292123c36538147d09f05bcdc843eb5572666a338d

Observation 5afcdcbc-5d7e-4568-a5f8-eedf8c1efd40 · outbound

This paper cites Channels Last Memory Format in PyTorch.

Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model Channels Last Memory Format in PyTorch

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T08:02:24.507443Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-19T08:02:23.002090Z digest=sha256:f60cc03c7a2f690f26a3681c323ed750c5350c55b167865e13a933872e401b58

Observation 96ebe783-8ae8-4e0a-9c07-6b88701ac625 · outbound

This paper cites Jordan, and Ion Stoica.

Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model Jordan, and Ion Stoica

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T08:02:24.511710Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-19T08:02:23.002090Z digest=sha256:506d0d30103e86640b92e6d8bd7278c975d1ed23fb98a74565367678618c3aa6

Observation 3b3631c5-b79b-4018-89aa-00fbec5f72e2 · outbound

This paper cites Pytorch rpc: Distributed deep learning built on tensor-optimized remote procedure calls.

Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model Pytorch rpc: Distributed deep learning built on tensor-optimized remote procedure calls

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T08:02:24.517347Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-19T08:02:23.002090Z digest=sha256:f5a6f8d4857795f8ef614293282462e2f25c6849717122d8b0d645009aedc8fd

Observation 1597fe2b-4566-4862-95e8-7fe26d69ce6c · outbound

This paper cites Mooncake: A KVCache-centric Disaggregated Architecture for LLM Serving.

Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model Mooncake: A KVCache-centric Disaggregated Architecture for LLM Serving

Reference 38

Resolution
verified exact
arxiv_id, observed 2026-05-19T08:02:23.720450Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-19T08:02:23.002090Z digest=sha256:831c48e6644592ddc94bdc50ad544c8d9795831840e5c4a30ca5af28b931dceb

Observation 33fa9489-9a01-4f62-aee0-54cce38f314f · outbound

This paper cites The llama 3 herd of models, April 2024.

Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model The llama 3 herd of models, April 2024

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T08:02:24.524622Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-19T08:02:23.002090Z digest=sha256:6e25a9f0a1d27131b2659134d646ddd8fb41d6ee62b9b88a58c7edae5e91d89c

Observation 6220c3a1-967c-444b-89f2-dd4e2e8bad5c · outbound

This paper cites PySceneDetect.

Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model PySceneDetect

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T08:02:24.529179Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-19T08:02:23.002090Z digest=sha256:83db6dc89e74f91719b1c0716a4d20bba48d9c7c504edb64685ac89de561a64c

Observation 4a718b9d-7051-47c9-abec-45dcd19d65de · outbound

This paper cites an unresolved cited work.

Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model Unresolved cited work

Reference 41

Resolution
unresolved
raw_fallback, observed 2026-05-19T08:02:24.533677Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-19T08:02:23.002090Z digest=sha256:169b108ce44bc0785226cb3d62c8f1bebf6ec7ef1206047a85a2721876fe8933

Observation 698c0a06-2230-4b30-a24f-5c333ae0832a · outbound

This paper cites Panda-70m: Captioning 70m videos with multiple cross-modality teachers.

Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model Panda-70m: Captioning 70m videos with multiple cross-modality teachers

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T08:02:24.538229Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-19T08:02:23.002090Z digest=sha256:6517cccb9e3f9469c163f5682868b5e79b218ebe7475c23c96a487e6ee40f4a3

Observation b69d111d-f1b2-40b2-9ff3-2723e9c50d31 · outbound

This paper cites Laion-5b: An open large-scale dataset for training next generation image-text models.

Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model Laion-5b: An open large-scale dataset for training next generation image-text models

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T08:02:24.544287Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-19T08:02:23.002090Z digest=sha256:424e560df6e79a5201eef13e0c04d0570cb43dcc3c55f6d3731d4f4ba7fb2708

Observation f3ed22ab-8066-4c94-948f-362cbeeefc68 · outbound

This paper cites Clip-based nsfw detector.

Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model Clip-based nsfw detector

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T08:02:24.548641Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-19T08:02:23.002090Z digest=sha256:331aa50d8323dd0eaaded8bde0d4add275f13e2f1b4188f8318c953e77f2d976

Observation 7de259bf-3738-4e5d-b8c2-a5445dc425c0 · outbound

This paper cites Efficientnet: Rethinking model scaling for convolutional neural networks.

Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model Efficientnet: Rethinking model scaling for convolutional neural networks

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T08:02:24.553407Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-19T08:02:23.002090Z digest=sha256:69e954be7bcbc5b5c6ef6087017d55ac3e8375cb9b743b945f5f39fc75472a30

Observation a364f9ec-d3d8-4035-b293-6d5c6951667c · outbound

This paper cites Paddleocr.

Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model Paddleocr

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T08:02:24.558593Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-19T08:02:23.002090Z digest=sha256:cc3e4edcc2dbdf5f4d34dd92dd3d1ff78ce8a19ef8a80bf0915fb2c8fcc1d2cf

Observation 7daad267-6ed3-44a5-b876-b5cda10d7c8d · outbound

This paper cites an unresolved cited work.

Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model Unresolved cited work

Reference 47

Resolution
unresolved
raw_fallback, observed 2026-05-19T08:03:01.620226Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-19T08:02:23.002090Z digest=sha256:a770de4c4acd6b2a5de86992d067faf934e1dcfdab3caf55e6eaed0fa302ae09

Observation 81735cd3-50dd-4f7c-aeb3-0f7bfcfd589a · outbound

This paper cites Diatom autofocusing in brightfield microscopy: a comparative study.

Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model Diatom autofocusing in brightfield microscopy: a comparative study

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T08:03:01.325066Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-19T08:02:23.002090Z digest=sha256:2a5c0d5a93067c56393193bbc6628bc32240e9189d282bf7fd69b0e68deb6f17

Observation 4ff1bcaa-a0ab-43ad-8a6c-765d98a938c5 · outbound

This paper cites Improving image generation with better captions.

Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model Improving image generation with better captions

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T08:03:01.467523Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-19T08:02:23.002090Z digest=sha256:fec825b8dc2b2b9c541dc26ff3e5d5cf2f1d3ebc3c39989e7aeed538bac11220

Observation 8ebdd587-9a73-4601-9a99-2bb5b964648d · outbound

This paper cites Some methods for classification and analysis of multivariate observations.

Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model Some methods for classification and analysis of multivariate observations

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T08:03:01.483371Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-19T08:02:23.002090Z digest=sha256:e2596b46eeb08ba87553fd68b57da50137910986d5876ab54364e50d91faedee

Observation f61c2c9b-4f03-49ce-a065-2a4b73c8d210 · outbound

This paper cites DSV: Exploiting Dynamic Sparsity to Accelerate Large-Scale Video DiT Training.

Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model DSV: Exploiting Dynamic Sparsity to Accelerate Large-Scale Video DiT Training

Reference 53

Resolution
verified exact
arxiv_id, observed 2026-05-19T08:02:23.330158Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-19T08:02:23.002090Z digest=sha256:5b8cd59910d194070bbf819b467010715cdfc573faa6ec5fffea7b4bc055207a

Observation 34702d07-6542-437e-a503-2fe9521d526b · outbound

This paper cites Diffusion Forcing: Next-token Prediction Meets Full-Sequence Diffusion.

Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model Diffusion Forcing: Next-token Prediction Meets Full-Sequence Diffusion

Reference 54

Resolution
metadata mismatch
arxiv_id, observed 2026-05-19T08:02:23.338415Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-19T08:02:23.002090Z digest=sha256:1ff7ccade977dc596c9a3b6fc2c08da8c2c2f33000b815f2cd08413d5fb9833f

Observation afda7481-5ba0-4838-b530-1711a28a1ee6 · outbound

This paper cites LTX-Video: Realtime Video Latent Diffusion.

Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model LTX-Video: Realtime Video Latent Diffusion

Reference 55

Resolution
verified exact
local_arxiv, observed 2026-05-19T08:02:23.451641Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-19T08:02:23.002090Z digest=sha256:bb4ab7d837ab07a93cb7ca62b3425e51b3dc5cb65aee50b509b760b0b8b57348

Observation 345c7c81-7ca0-4f60-9dcd-030355ebec85 · outbound

This paper cites Taming Teacher Forcing for Masked Autoregressive Video Generation.

Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model Taming Teacher Forcing for Masked Autoregressive Video Generation

Reference 56

Resolution
verified exact
arxiv_id, observed 2026-05-19T08:02:23.524350Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-19T08:02:23.002090Z digest=sha256:86bd54c22508dbd9c62a1a8157d6805aab41cd3cb5a7c95e2e141526fd928140

Observation f6b0c620-dc27-4df5-b667-8f25f95a4d3e · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 57

Resolution
verified exact
local_arxiv, observed 2026-05-19T08:02:23.544061Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-19T08:02:23.002090Z digest=sha256:395128f83baf8b108203d01f9c20d8d5b2b1e09d25229604f1666e12dbba14fd

Observation 383236cd-3173-4e19-a441-9fb768e7a307 · outbound

This paper cites 2024 , journal =.

Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model 2024 , journal =

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T08:03:01.423896Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-19T08:02:23.002090Z digest=sha256:d9cccfe47b730416e9e7f0ba65734e1788bb57194449b51b8a7c3ffc2cf37724

Observation 6bb8a284-d037-4930-b7bc-82170cd6f73a · outbound

This paper cites 2024 , eprint=.

Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model 2024 , eprint=

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T08:03:01.500827Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-19T08:02:23.002090Z digest=sha256:8f5bf3c5ea42d271e69a09bbf431a4eec40f9bf14f1f137941b239ca3267d92d

Observation 1e4c5300-0089-4db7-aaf8-9dda2e70d401 · outbound

This paper cites 2024 , eprint=.

Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model 2024 , eprint=

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T08:03:01.356495Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-19T08:02:23.002090Z digest=sha256:588498b6cdf522c765066cdb19b44f59ab5d17b48795b223f87c4aa724af7641

Observation b2d931d3-ba4b-4117-8aab-26e23e1a1725 · outbound

This paper cites 2025 , eprint=.

Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model 2025 , eprint=

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T08:03:01.606674Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-19T08:02:23.002090Z digest=sha256:734890470422944f9957b9a60aafca36ad7e99ac8308ad2a5a592329fc8cc000

Observation 56820fb5-a3f8-4e25-8b24-51a84497a81d · outbound

This paper cites Computer Science.

Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model Computer Science

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T08:03:01.561945Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-19T08:02:23.002090Z digest=sha256:0530862ce44b652243d9593b6544df1465dbd758eb1f01d86fa70168d61b760c

Observation 04445759-84d1-4007-a6b6-aa4f445b0e57 · outbound

This paper cites 2024 , eprint=.

Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model 2024 , eprint=

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T08:03:01.280640Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-19T08:02:23.002090Z digest=sha256:021d5ab908c7a4a72d503d0d9579042f71e44eb210ed5e90114e086ae20a7800

Observation 9e582ab3-6d34-4dbf-b892-555e7ecb7c85 · outbound

This paper cites 2023 , eprint=.

Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model 2023 , eprint=

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T08:03:01.583393Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-19T08:02:23.002090Z digest=sha256:7a581eece22a63fa0203612c020ebb01facbca749715ce98098585d218fd5957

Observation ec80df0b-057d-40ac-bd55-801dabc53a30 · outbound

This paper cites 2024 , eprint=.

Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model 2024 , eprint=

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T08:03:01.618265Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-19T08:02:23.002090Z digest=sha256:2fdd0cf94eba3452b71e61658b65e7dfa3a728a892b8a6af363bdedbad60bcb1

Observation 5ee362d6-1eff-4af4-a14c-8976d9597019 · outbound

This paper cites 2024 , eprint=.

Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model 2024 , eprint=

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T08:03:01.337535Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-19T08:02:23.002090Z digest=sha256:d35307ebf75b513d30c221b8fa063350aa34c80d9c84888794b7ed27d32c0080

Observation 938b2f6f-f49b-42f9-a20e-594e48df9a9c · outbound

This paper cites CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer.

Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer

Reference 67

Resolution
metadata mismatch
local_arxiv, observed 2026-05-19T08:02:23.981187Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-19T08:02:23.002090Z digest=sha256:2b59d410465e7046b6cb319c91a63cc2f0149772f90ab3330bd6ff7757265942

Observation 208833bf-b0ca-4eef-a99c-11d0bc2de12a · outbound

This paper cites 2024 , url =.

Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model 2024 , url =

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T08:03:01.546048Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-19T08:02:23.002090Z digest=sha256:e73f703868fbe6156e3e50e25cbfb980bcddf9a9a0460401092a91f051294b3a

Observation 7475e662-ce52-48f3-97e8-f00428f6dfd3 · outbound

This paper cites Open-Sora Plan: Open-Source Large Video Generation Model.

Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model Open-Sora Plan: Open-Source Large Video Generation Model

Reference 69

Resolution
metadata mismatch
local_arxiv, observed 2026-05-19T08:02:23.993679Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-19T08:02:23.002090Z digest=sha256:4d1ab15b46de8585c187bdc0e29927b1c26f4d5f8b3276b0f0993c5d9604569c

Observation c90974cd-bf48-4d76-9ec8-ca9582573bcc · outbound

This paper cites 2025 , eprint=.

Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model 2025 , eprint=

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T08:03:01.573628Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-19T08:02:23.002090Z digest=sha256:38851fa95d71916d8240364eb9e694b6174c64a6be49fbd934bee48853bbd798

Observation 9aea7952-175a-49e6-bf36-687b94cbb0c9 · outbound

This paper cites 2024 , journal =.

Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model 2024 , journal =

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T08:03:01.463370Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-19T08:02:23.002090Z digest=sha256:39020fa3457001b0f38e4bf51ed1cf86fc2daea57f4472e6f450b6f2b7a37981

Observation bde6ba1a-2c12-43db-a679-a77e55f9fb45 · outbound

This paper cites 2024 , journal =.

Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model 2024 , journal =

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T08:03:01.553871Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-19T08:02:23.002090Z digest=sha256:d2d50198475e9fc4083bd3bb68261a67a4b4b1921a8d0fec0d763781593cb7ca

Observation 9e817465-1af0-458a-8e07-2ac2c852c9de · outbound

This paper cites 2024 , journal =.

Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model 2024 , journal =

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T08:03:01.592156Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-19T08:02:23.002090Z digest=sha256:2a2a3d6d1464fff89583c53524b0c189c46b2bb7232b1fb8efed5419e7d6a7ab

Observation c86d4304-38f4-4c2a-8e98-6c088b0c8ca9 · outbound

This paper cites 2024 , journal =.

Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model 2024 , journal =

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T08:03:01.602775Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-19T08:02:23.002090Z digest=sha256:d8e8ea2a62813e188f15e778e9bf7292b88b1248ddb27a4d5ab9464d8dad88c3

Observation ed2c5a0d-8d5a-4d85-948a-e708b64277b1 · outbound

This paper cites 2025 , eprint=.

Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model 2025 , eprint=

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T08:03:01.351674Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-19T08:02:23.002090Z digest=sha256:2b17a2ba9b4933c94cf75fc89c187925b1bf88c8554b7c4a2e39f99254b18021

Observation 5cfb6c52-bd07-458c-970a-895d3de02a3e · outbound

This paper cites 2023 , eprint=.

Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model 2023 , eprint=

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T08:03:01.504512Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-19T08:02:23.002090Z digest=sha256:843cea56c3d7196b974eb1d4352e5ca43e90f71bbbf17161c531433b94578ab7

Observation 582de9a2-1b1b-41c2-a53b-f398898b442e · outbound

This paper cites and Ilharco, Gabriel and Song, Shuran and Kollar, Thomas and Carmon, Yair and Dave, Achal and Heckel, Reinhard and Muennighoff, Niklas and Schmidt, Ludwig , title=.

Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model and Ilharco, Gabriel and Song, Shuran and Kollar, Thomas and Carmon, Yair and Dave, Achal and Heckel, Reinhard and Muennighoff, Niklas and Schmidt, Ludwig , title=

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T08:03:01.528514Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-19T08:02:23.002090Z digest=sha256:ad7cdfa72ded5ff67837c4bdff32b6dbf68a9642f2bf48fde2b34730ace41fb9

Observation 5169f380-c5e0-4def-bdd1-48a43e8457ca · outbound

This paper cites 2023 , eprint=.

Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model 2023 , eprint=

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T08:03:01.530394Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-19T08:02:23.002090Z digest=sha256:380ad4c451280b011122cc637fc01ad9164a6651d221e32da686c22b2d123922

Observation 05f79cd2-3e60-472a-a480-73b736354cae · outbound

This paper cites PaLM 2 Technical Report.

Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model PaLM 2 Technical Report

Reference 79

Resolution
metadata mismatch
local_arxiv, observed 2026-05-19T08:02:23.360355Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-19T08:02:23.002090Z digest=sha256:aabfa120a42dde94b18960e5a3d0dc1a88a3d816bf4debf1fd688525e7db86de

Observation c6ceac81-6d16-4aeb-98e0-477ea8bbe813 · outbound

This paper cites CCNet: Extracting High Quality Monolingual Datasets from Web Crawl Data.

Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model CCNet: Extracting High Quality Monolingual Datasets from Web Crawl Data

Reference 80

Resolution
metadata mismatch
arxiv_id, observed 2026-05-19T08:02:23.384376Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-19T08:02:23.002090Z digest=sha256:0bdf71d5c093df170eb9adf87549c252fbc581ae153aa247f020c7adc7ffd349

Observation f66bfdee-d959-4ca7-9f0a-fdb399b09517 · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model LLaMA: Open and Efficient Foundation Language Models

Reference 81

Resolution
metadata mismatch
local_arxiv, observed 2026-05-19T08:02:23.391752Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-19T08:02:23.002090Z digest=sha256:08a6feff6686dc5eacb9e05de0d883403a1d9ef9c611184ad6c6cbdcb7821dc1

Observation 9a4bdb57-52db-44ba-84ea-5b98a6f043b1 · outbound

This paper cites Mistral 7B.

Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model Mistral 7B

Reference 82

Resolution
metadata mismatch
local_arxiv, observed 2026-05-19T08:02:23.415092Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-19T08:02:23.002090Z digest=sha256:1f3cdfd60e3b5e083bbda27dfd5a9453eef4ff26bb4559e13a8646501da60a72

Observation cf5eedea-4314-42f7-bb11-311f44ac5782 · outbound

This paper cites The Twelfth International Conference on Learning Representations , year=.

Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model The Twelfth International Conference on Learning Representations , year=

Reference 83

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T08:03:01.532515Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-19T08:02:23.002090Z digest=sha256:4df8060bfaf25e57a39b2f2311c8b65524e7fc500b428c28d2a5a88d9633c608

Observation 42719aa2-0294-4d5b-8abf-a17cc92315af · outbound

This paper cites an unresolved cited work.

Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model Unresolved cited work

Reference 84

Resolution
unresolved
raw_fallback, observed 2026-05-19T08:03:01.459534Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-19T08:02:23.002090Z digest=sha256:51dbb16a91eb54016964aa178f7df904e8fd816fdd2887f0ee6e1add458a9be8

Observation 52617cc4-a03e-4774-afd0-4e755f00c81d · outbound

This paper cites StarCoder: may the source be with you! , journal =.

Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model StarCoder: may the source be with you! , journal =

Reference 85

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T08:03:01.561011Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-19T08:02:23.002090Z digest=sha256:1c25247616ed0400def78c5f0cfc44f61dac050d86abc41b588b48e537e05e0f

Observation 80514e52-fa1d-47df-a965-6393ddb5fc42 · outbound

This paper cites 2024 , eprint=.

Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model 2024 , eprint=

Reference 86

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T08:03:01.512343Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-19T08:02:23.002090Z digest=sha256:33b962f6968471f614ea78cc83c85c61dca6e401799201140abcd9d4c84c004d

Observation 2e8c8a6a-7384-4052-9e62-380090c3f296 · outbound

This paper cites Gemma: Open Models Based on Gemini Research and Technology.

Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model Gemma: Open Models Based on Gemini Research and Technology

Reference 87

Resolution
metadata mismatch
local_arxiv, observed 2026-05-19T08:02:23.571195Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-19T08:02:23.002090Z digest=sha256:6d24baa556a5de1d3d7417c5c46be12ec362993fba52ef3970848c6d28887ee9

Observation 20f033f6-595f-4f2b-9f65-8e88f1537a10 · outbound

This paper cites Qwen Technical Report.

Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model Qwen Technical Report

Reference 88

Resolution
metadata mismatch
local_arxiv, observed 2026-05-19T08:02:23.578432Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-19T08:02:23.002090Z digest=sha256:4209ae7c11c8ef20d36d5404c08ce125a9bf91a32cd9425a7424798ae72c6a08

Observation f8a02304-ef88-4a23-a199-55336e0b03dd · outbound

This paper cites an unresolved cited work.

Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model Unresolved cited work

Reference 89

Resolution
unresolved
raw_fallback, observed 2026-05-19T08:03:01.470987Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-19T08:02:23.002090Z digest=sha256:9e853f8a0b9f158d72e9e86f59e6b2253f62559b132077b9c676b2b969ed78d0

Observation c1a9ab32-f2c9-404c-a30e-e0f51168c90c · outbound

This paper cites Microsoft Research Blog , year=.

Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model Microsoft Research Blog , year=

Reference 90

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T08:03:01.556415Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-19T08:02:23.002090Z digest=sha256:be241cb22a095104323f7ec220218183d543c632d216e9814a2e48d07af8547c

Observation 375cea56-7ee4-46c8-8e7d-f8386e8f1a49 · outbound

This paper cites DeepSeek LLM: Scaling Open-Source Language Models with Longtermism.

Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model DeepSeek LLM: Scaling Open-Source Language Models with Longtermism

Reference 91

Resolution
metadata mismatch
local_arxiv, observed 2026-05-19T08:02:23.605506Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-19T08:02:23.002090Z digest=sha256:dae5ab2ff60548d6a2f795985bb8fa9f1377ce8bf9a3b5a4bbc5b70d0be4979c

Observation 2c5cd385-a1f4-4238-929f-7be6eae0d30b · outbound

This paper cites Code Llama: Open Foundation Models for Code.

Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model Code Llama: Open Foundation Models for Code

Reference 92

Resolution
metadata mismatch
local_arxiv, observed 2026-05-19T08:02:23.611249Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-19T08:02:23.002090Z digest=sha256:3b61d2e1ae8aac101a60ef34c964e0616176be4ad0ff5e4acaab4c9349bc3f12

Observation d37b6d8b-f2e7-4ca1-8e66-2b13acfb6b95 · outbound

This paper cites Advances in Neural Information Processing Systems , volume=.

Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model Advances in Neural Information Processing Systems , volume=

Reference 93

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T08:03:01.536557Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-19T08:02:23.002090Z digest=sha256:c3e9bd84b2e41a964c9e281e47b73d23e4dc7476add66d139abf94b7a11ebd41

Observation 4f79123f-43bb-42fb-a6e8-7b6ada5e23df · outbound

This paper cites Llemma: An Open Language Model For Mathematics.

Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model Llemma: An Open Language Model For Mathematics

Reference 94

Resolution
metadata mismatch
arxiv_id, observed 2026-05-19T08:17:47.048289Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-19T08:02:23.002090Z digest=sha256:bd0dbd4c9512c9d8a70da1c308f2d56bb97781429c72babb4da84bcf87f4d5ef

Observation 3179ab27-421d-4ecb-b699-ce7525ad7db8 · outbound

This paper cites InternLM-Math: Open Math Large Language Models Toward Verifiable Reasoning.

Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model InternLM-Math: Open Math Large Language Models Toward Verifiable Reasoning

Reference 95

Resolution
metadata mismatch
arxiv_id, observed 2026-05-19T08:02:23.699532Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-19T08:02:23.002090Z digest=sha256:12cb46197bd0f9920af547d7b2df9f50acc3542edf7fbc8cdbce00c1487ea8d0

Observation d4b983ec-f9d8-4e09-9f41-7173e8d0315d · outbound

This paper cites Let's Verify Step by Step.

Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model Let's Verify Step by Step

Reference 96

Resolution
metadata mismatch
local_arxiv, observed 2026-05-19T08:02:23.712669Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-19T08:02:23.002090Z digest=sha256:2f4531fa637b12530cfe6e8df614adbe6b0619d5680791d725fc79a94792f4c3

Observation ed18654e-3653-4deb-b6da-9f7ac6f729de · outbound

This paper cites The Twelfth International Conference on Learning Representations , year=.

Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model The Twelfth International Conference on Learning Representations , year=

Reference 97

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T08:03:01.419752Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-19T08:02:23.002090Z digest=sha256:9996025686bedfdb43a9dead798d5bf076941a40b65430a797fcd73a9aa0b228

Observation 9f330a17-c10e-487b-8698-8375c9dbe106 · outbound

This paper cites Large-scale Dataset Pruning with Dynamic Uncertainty.

Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model Large-scale Dataset Pruning with Dynamic Uncertainty

Reference 98

Resolution
verified exact
arxiv_id, observed 2026-05-19T08:02:23.727730Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-19T08:02:23.002090Z digest=sha256:6a3c0f1e265a46f572e6401abea86641909430a1047efd235a5bfca599c725e8

Observation d345aa5e-0f29-4cda-83f2-4a05c6132b5c · outbound

This paper cites NIPS , volume=.

Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model NIPS , volume=

Reference 99

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T08:03:01.438254Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-19T08:02:23.002090Z digest=sha256:6d7bb4641f6644de00e0d5e9c42fe3b6167efaa6e300edf491fc8a27246746a8

Observation 755a65f1-c9ee-420e-b1de-c12ed84ad7d4 · outbound

This paper cites From Quantity to Quality: Boosting LLM Performance with Self-Guided Data Selection for Instruction Tuning.

Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model From Quantity to Quality: Boosting LLM Performance with Self-Guided Data Selection for Instruction Tuning

Reference 100

Resolution
verified exact
arxiv_id, observed 2026-05-19T08:02:23.882285Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-19T08:02:23.002090Z digest=sha256:1d6e9a7ba896b1d4a01f3c019e9b5c6b6d998f4598a133d04eb53aacdfa0eeb0

Observation e19c190a-a269-4b39-bc60-e79be703cff5 · outbound

This paper cites One-Shot Learning as Instruction Data Prospector for Large Language Models.

Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model One-Shot Learning as Instruction Data Prospector for Large Language Models

Reference 101

Resolution
verified exact
arxiv_id, observed 2026-05-19T08:02:23.902576Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-19T08:02:23.002090Z digest=sha256:3985026f0d78c2e90d7451470b28d21f0e778983f14bc296befaf5c278c19e2b

Observation 0b3012be-c414-4d6c-81bb-f3c4215a930a · outbound

This paper cites Self-Evolved Diverse Data Sampling for Efficient Instruction Tuning.

Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model Self-Evolved Diverse Data Sampling for Efficient Instruction Tuning

Reference 102

Resolution
metadata mismatch
arxiv_id, observed 2026-05-19T08:02:23.909152Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-19T08:02:23.002090Z digest=sha256:0b224db7c10fef97b770f5c354e7aad78e9191ae302d8dddbb8f83e9b531c223

Observation 510cf5ae-db2d-4c7d-a76b-282746e4a2eb · outbound

This paper cites ICLR , year=.

Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model ICLR , year=

Reference 103

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T08:03:01.599136Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-19T08:02:23.002090Z digest=sha256:7e270c2439cea70da46754d7cae4255c384e86d456533159d12ec270b6a07030

Observation 85d0480e-642e-47a8-86f7-1f6c1e09f8b1 · outbound

This paper cites ICLR , year=.

Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model ICLR , year=

Reference 104

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T08:03:01.600880Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-19T08:02:23.002090Z digest=sha256:40b351b1989117f8ea988222f32024bd25419fe5aa2a8d0d4a166eb4518edbbf

Pith citing papers

Observation 44c1afa1-920b-4ea5-98cf-fd32047b824f · inbound

Wan: Open and Advanced Large-Scale Video Generative Models cites this paper.

Wan: Open and Advanced Large-Scale Video Generative Models Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model

Reference 33

Resolution
verified exact
local_arxiv, observed 2026-05-22T23:07:14.457899Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-22T23:05:32.595632Z digest=sha256:b21cd8b04fecda7a06386ce0a403646670e811d12b7f0b29c1998f7607e6575a

Observation 31dcae6c-86f1-42e0-880c-09d71cdcfaf0 · inbound

VBench-2.0: Advancing Video Generation Benchmark Suite for Intrinsic Faithfulness cites this paper.

VBench-2.0: Advancing Video Generation Benchmark Suite for Intrinsic Faithfulness Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model

Reference 62

Resolution
verified exact
arxiv_id, observed 2026-05-19T08:02:24.560574Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-14T18:42:02.940250Z digest=sha256:9787c60fb3bb59e161e4b6ba45172e4a571689db036584f842175bd6ee9c56ce

Observation eed6a369-57ed-43b2-a4a0-53823b0bb2fe · inbound

MAGI-1: Autoregressive Video Generation at Scale cites this paper.

MAGI-1: Autoregressive Video Generation at Scale Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-05-19T08:02:24.560574Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-13T20:31:15.700943Z digest=sha256:6f513af10c25721eeca7eaa95a6ec758f693785c2b801b1e42905f76e3fda178

Observation 0d4a097e-ebf6-431b-9307-90e7333825a0 · inbound

GenHSI: Controllable Generation of Human-Scene Interaction Videos cites this paper.

GenHSI: Controllable Generation of Human-Scene Interaction Videos Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model

Reference 61

Resolution
verified exact
arxiv_id, observed 2026-05-19T08:02:24.560574Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T07:25:41.219753Z digest=sha256:7a6e83bb6aeef2307af920bb78a1bdf56c5c9737fa625d1ffbf81b24700d38ab

Observation 6b1207d8-4f4c-4907-b48b-14a385e92439 · inbound

Listener-Rewarded Thinking in VLMs for Image Preferences cites this paper.

Listener-Rewarded Thinking in VLMs for Image Preferences Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-19T08:02:24.560574Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T07:38:52.273903Z digest=sha256:d3aa6073c527e6daba51b15030a3d35b7ad77046f98681a6fde40374a91da13c

Observation 8802044d-ba60-45f8-967f-f8ae9eaab0df · inbound

Waver: Wave Your Way to Lifelike Video Generation cites this paper.

Waver: Wave Your Way to Lifelike Video Generation Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-05T17:45:17.186768Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T17:45:17.186768Z digest=sha256:74fa0d4ad32c92ac06bd093d8bc08039ea8b64a2c633317d272649c8cda472b8

Observation 566e6d69-9a7f-4b12-a148-f183e1019dc6 · inbound

UniVerse-1: Unified Audio-Video Generation via Stitching of Experts cites this paper.

UniVerse-1: Unified Audio-Video Generation via Stitching of Experts Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-05T00:05:47.622882Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T00:05:47.622882Z digest=sha256:e1ad03282330cc01b169da9f5ef5435acb612dfbc5aa3d82551752e7e262c966

Observation e5445f7e-d9e2-48ea-89bf-662cbd1b64c2 · inbound

RewardDance: Reward Scaling in Visual Generation cites this paper.

RewardDance: Reward Scaling in Visual Generation Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-04T20:08:58.832391Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:08:58.832391Z digest=sha256:a7ebaad714443b3f1f668266b91ec2a0614928ece96ab104374753c43ae72e45

Observation 70974366-1796-44bb-9adf-73557c0f3467 · inbound

Improving Video Diffusion Transformer Training by Multi-Feature Fusion and Alignment from Self-Supervised Vision Encoders cites this paper.

Improving Video Diffusion Transformer Training by Multi-Feature Fusion and Alignment from Self-Supervised Vision Encoders Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-04T18:58:10.953460Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:58:10.953460Z digest=sha256:7261afc6617bfa632e9a4b2b33000d97cd22fdb372389ecab5e5b5824518509f

Observation b966e68f-0b2e-4e8e-af02-c24edfaacfdd · inbound

Enhancing Physical Plausibility in Video Generation by Reasoning the Implausibility cites this paper.

Enhancing Physical Plausibility in Video Generation by Reasoning the Implausibility Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-19T08:02:24.560574Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-18T12:55:42.679016Z digest=sha256:2a2d03b4e0e980df40e99464747ebf1d89bd3a06e807ee5f98003efe874ca768

Observation 8a4bf5bb-9ddc-4ecb-9c1e-0b0b518987e1 · inbound

UniVideo: Unified Understanding, Generation, and Editing for Videos cites this paper.

UniVideo: Unified Understanding, Generation, and Editing for Videos Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-04T10:49:48.658757Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T10:49:48.658757Z digest=sha256:941db68f241f09830a1c04437810e1c60da235a8186b27273ea6955dc7248448

Observation e52145d8-551f-44ff-963f-decce4ea2148 · inbound

HunyuanVideo 1.5 Technical Report cites this paper.

HunyuanVideo 1.5 Technical Report Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-19T08:02:24.560574Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T02:31:34.927101Z digest=sha256:85574980669dba6c39c974fe8067abf0cf92ba4638785a8488d732ee9c612e55

Observation 78382ff9-3617-408b-b8a7-f2b23c649fe4 · inbound

Delving into Latent Spectral Biasing of Video VAEs for Superior Diffusability cites this paper.

Delving into Latent Spectral Biasing of Video VAEs for Superior Diffusability Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-03T18:25:55.724129Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T18:25:55.724129Z digest=sha256:f3daa7ef22c490653011212a7861ab91e0a60b1a0a8f96f422014a00d2b48ac5

Observation 58b1a2bb-31ae-4566-bf77-29274feb2d93 · inbound

VideoASMR-Bench: Can AI-Generated ASMR Videos Fool VLMs and Humans? cites this paper.

VideoASMR-Bench: Can AI-Generated ASMR Videos Fool VLMs and Humans? Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-05-19T08:02:24.560574Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T21:46:43.353305Z digest=sha256:66d3cc531c95aa2299987b960ef507e0563c183f573a9fdb688936beeb4e4b70

Observation 98638ac3-139f-431d-9cf9-f47b4921b9ca · inbound

Self-transcendence: Is External Feature Guidance Indispensable for Accelerating Diffusion Transformer Training? cites this paper.

Self-transcendence: Is External Feature Guidance Indispensable for Accelerating Diffusion Transformer Training? Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-03T11:04:32.252418Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T11:04:32.252418Z digest=sha256:f5fd55e63711bdc520cb6d1eb84aa607ab7ea30ed25991c853ac3e452674955b

Observation 6c7e2aae-f7a7-4913-91ca-698675b9abf6 · inbound

Beyond Rigid: Benchmarking Non-Rigid Video Editing cites this paper.

Beyond Rigid: Benchmarking Non-Rigid Video Editing Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model

Reference 379

Resolution
unresolved
no resolver link, observed 2026-08-03T08:04:51.652932Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T08:04:51.652932Z digest=sha256:9418f4d5e5bb133f5ab2a0fe758bd82b5ff1e2fe81954b003a54371b7c27658d

Observation 9edbeec9-48e9-47c3-9ad7-030279c654b5 · inbound

CG-MLLM: Captioning and Generating 3D content via Multi-modal Large Language Models cites this paper.

CG-MLLM: Captioning and Generating 3D content via Multi-modal Large Language Models Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model

Reference 37

Resolution
metadata mismatch
local_arxiv, observed 2026-05-21T14:50:14.801537Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-21T14:48:21.787919Z digest=sha256:8ad28ef555d6d198e88c5ed7c8df5b965b92b87642b9fccd895f5c6481b445d4

Observation 9deb6519-0c9e-448f-bdd1-1f7fa87fdf40 · inbound

SynthForensics: Benchmarking and Evaluating People-Centric Synthetic Video Deepfakes cites this paper.

SynthForensics: Benchmarking and Evaluating People-Centric Synthetic Video Deepfakes Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model

Reference 38

Resolution
verified exact
arxiv_id, observed 2026-05-19T08:02:24.560574Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T07:32:30.580643Z digest=sha256:208def05536288a9eb50d5f14ee3ff13ced554441e024ec99526a40e34ad3aac

Observation 9c7ac717-3e3a-48fb-9ebb-72ff1e99671f · inbound

EchoTorrent: Towards Swift, Sustained, and Streaming Multi-Modal Video Generation cites this paper.

EchoTorrent: Towards Swift, Sustained, and Streaming Multi-Modal Video Generation Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model

Reference 60

Resolution
verified exact
arxiv_id, observed 2026-05-19T08:02:24.560574Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T22:20:16.320171Z digest=sha256:173d0980a32653ca636e7905b52297e804641ffc1f6a6170470aa030e3538b04

Observation 5112e65a-e406-4eb6-865b-585573014806 · inbound

Event-Driven Video Generation cites this paper.

Event-Driven Video Generation Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model

Reference 17

Resolution
unresolved
no resolver link, observed 2026-07-14T22:56:23.337858Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T22:56:23.337858Z digest=sha256:872fe6464cd1fe1647386e94ec44c04f87c78b97f005f0ec6b59502c7991ea13

Observation 11af68b9-8d46-46fb-895f-0186e44043c4 · inbound

ActionParty: Multi-Subject Action Binding in Generative Video Games cites this paper.

ActionParty: Multi-Subject Action Binding in Generative Video Games Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-03T00:56:28.335209Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T00:56:28.335209Z digest=sha256:67519daf47df4202895cbf0ed038f31fec997be47e1f72d54536af0079a131c2

Observation 45280eac-eeb0-46f0-856c-4213cafa5701 · inbound

Evolution of Video Generative Foundations cites this paper.

Evolution of Video Generative Foundations Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model

Reference 82

Resolution
verified exact
arxiv_id, observed 2026-05-19T08:02:24.560574Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T18:41:38.616611Z digest=sha256:c3c65dcd60f7b35b75b5ccee51b991c1de2869b0364acd857bdf5a38fbff1112

Observation a0c1f447-6ba6-4461-a89f-a9e5ca911a87 · inbound

Efficient Video Diffusion Models: Advancements and Challenges cites this paper.

Efficient Video Diffusion Models: Advancements and Challenges Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model

Reference 100

Resolution
verified exact
arxiv_id, observed 2026-05-19T08:02:24.560574Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T08:28:29.706249Z digest=sha256:e9c154bd424d6ad9b75a010b6f06ac52951c17e0026a21afd92dd55683d24550

Observation 6616cf1e-aa88-4c11-b187-2035ae3fb2cb · inbound

Motif-Video 2B: Technical Report cites this paper.

Motif-Video 2B: Technical Report Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-19T08:02:24.560574Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T15:50:04.001693Z digest=sha256:b8abe4ab446f8f9b2586af37fa25987aada7f299dd8864b8fe4ea2bb27126de6

Observation 5e529286-57be-4923-bc51-63bd9064b1ce · inbound

Motif-Video 2B: Technical Report cites this paper.

Motif-Video 2B: Technical Report Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model

Reference 24

Resolution
verified exact
local_arxiv, observed 2026-05-21T00:13:53.097527Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-21T00:12:23.096146Z digest=sha256:9ac12f321136ab1a2403c3cb777d39bd82dec6ee6a9d7a2c1b7e0e68830c6e73

Observation ce6aff6d-cd09-4797-8266-a6ec2589d1e2 · inbound

DreamShot: Personalized Storyboard Synthesis with Video Diffusion Prior cites this paper.

DreamShot: Personalized Storyboard Synthesis with Video Diffusion Prior Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-19T08:02:24.560574Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T07:21:45.633310Z digest=sha256:1d5c24bfef2f817dac5610ae4b4e3797cc599919079f59a969405d8b4266ef91

Observation f55539d9-13c8-43c1-9ab3-87e8da5cf939 · inbound

Leveraging Verifier-Based Reinforcement Learning in Image Editing cites this paper.

Leveraging Verifier-Based Reinforcement Learning in Image Editing Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model

Reference 38

Resolution
verified exact
arxiv_id, observed 2026-05-19T08:02:24.560574Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-07T08:00:33.307429Z digest=sha256:879fb34254e1e1ec32159c29315e24090dc61ffb1965558993d755ad5a07aafa

Observation 87934217-edb9-465f-930d-db6142090418 · inbound

Leveraging Verifier-Based Reinforcement Learning in Image Editing cites this paper.

Leveraging Verifier-Based Reinforcement Learning in Image Editing Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model

Reference 38

Resolution
verified exact
local_arxiv, observed 2026-05-21T09:14:05.928718Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-21T09:11:02.183133Z digest=sha256:3ca26a83fefc0f734e414f79232bff085b9ba358aa3753c2ed72e57a8ab0d88c

Observation 0ce52856-1c69-46a6-804b-832e1d875b22 · inbound

Offline Preference Optimization for Rectified Flow with Noise-Tracked Pairs cites this paper.

Offline Preference Optimization for Rectified Flow with Noise-Tracked Pairs Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-19T08:02:24.560574Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T02:10:27.595446Z digest=sha256:2063b69be03055902e174f82026839a2e9bd722497430149d1c3798c3d10fbb0

Observation 84501807-5c5d-40ce-ab05-21253e441eaf · inbound

Qwen-Image-2.0 Technical Report cites this paper.

Qwen-Image-2.0 Technical Report Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-19T08:02:24.560574Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T04:21:18.312613Z digest=sha256:f0674837697aac082dde9f3df5671b4b21422bf437233119984fb35534908699

Observation 01c147af-d55c-4188-99ee-759ec01aca87 · inbound

HorizonDrive: Self-Corrective Autoregressive World Model for Long-horizon Driving Simulation cites this paper.

HorizonDrive: Self-Corrective Autoregressive World Model for Long-horizon Driving Simulation Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-19T08:02:24.560574Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-13T01:16:15.422226Z digest=sha256:5fa5ae615257aeb2b0e3d76422b55391084e06042ae486a7dcdbe4062e176be1

Observation d427552a-f2fe-4715-a1fb-3be5be616e75 · inbound

HorizonDrive: Self-Corrective Autoregressive World Model for Long-horizon Driving Simulation cites this paper.

HorizonDrive: Self-Corrective Autoregressive World Model for Long-horizon Driving Simulation Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model

Reference 14

Resolution
verified exact
local_arxiv, observed 2026-05-25T06:35:24.933805Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-25T06:31:26.421088Z digest=sha256:ebec83a3ec348ebc4bcf7807311489cfdaed7cf50dc9c24fe2d3aa20fcda2fd7

Observation 9a1d3f35-bb45-4a63-9302-acba56d47bde · inbound

Qwen-Image-VAE-2.0 Technical Report cites this paper.

Qwen-Image-VAE-2.0 Technical Report Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model

Reference 12

Resolution
metadata mismatch
arxiv_id, observed 2026-05-19T08:02:24.560574Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-14T19:53:25.849049Z digest=sha256:a6855f990ffe25edf485e1f4f63a21159f1c90063d95e24c536929b52ea3d3e3

Observation 658e9d79-8cae-437e-943c-f6038e0cc3e8 · inbound

HASTE: Training-Free Video Diffusion Acceleration via Head-Wise Adaptive Sparse Attention cites this paper.

HASTE: Training-Free Video Diffusion Acceleration via Head-Wise Adaptive Sparse Attention Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-19T08:02:24.560574Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T05:17:48.403406Z digest=sha256:29ef65056048c3996cd36a8023aaf0d5cbe09902ff65a5daf7c7ac70ddc511d0

Observation b95885bf-e0f8-4cda-a132-16357f26e81a · inbound

MechVerse: Evaluating Physical Motion Consistency in Video Generation Models cites this paper.

MechVerse: Evaluating Physical Motion Consistency in Video Generation Models Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model

Reference 25

Resolution
verified exact
local_arxiv, observed 2026-07-01T14:25:46.377582Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-30T21:37:15.414734Z digest=sha256:09ec163609f4436bcb6a9d3203089e2cd71da0cdbc44fafbc24c9d85750d7ad4

Observation bca9dbe6-be37-4f29-ad2b-7996904c7f08 · inbound

RAVEN: Real-time Autoregressive Video Extrapolation with Consistency-model GRPO cites this paper.

RAVEN: Real-time Autoregressive Video Extrapolation with Consistency-model GRPO Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-07-01T14:35:47.172418Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-30T20:38:28.328436Z digest=sha256:460c03b16e60c43a85c5b6b2514fd53c2da34f00ef342e818c26ccbe1c425915

Observation ecead28d-b9d4-4590-ab41-ecf0a83832a6 · inbound

AtlasVid: Efficient Ultra-High-Resolution Long Video Generation via Decoupled Global-Local Modeling cites this paper.

AtlasVid: Efficient Ultra-High-Resolution Long Video Generation via Decoupled Global-Local Modeling Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-05-20T18:28:53.009080Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-20T18:26:57.735246Z digest=sha256:8b181edb9d7df615eaf97c39307daa78d23abc459aaf09e0010b8a2fdccd47ca

Observation 58ef28cc-f69a-4fb4-82e9-fea7321a3ea3 · inbound

Lance: Unified Multimodal Modeling by Multi-Task Synergy cites this paper.

Lance: Unified Multimodal Modeling by Multi-Task Synergy Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model

Reference 79

Resolution
verified exact
local_arxiv, observed 2026-05-20T11:48:14.937697Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-20T11:46:52.658984Z digest=sha256:91e0e989feceb0d83ce03ff2f4e7f3649185578b4af3a5d1414eb034307db288

Observation 8e42cb5d-e659-4374-9e79-958876ff0e86 · inbound

Lance: Unified Multimodal Modeling by Multi-Task Synergy cites this paper.

Lance: Unified Multimodal Modeling by Multi-Task Synergy Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model

Reference 80

Resolution
verified exact
local_arxiv, observed 2026-05-21T07:59:50.631393Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-21T07:56:34.034047Z digest=sha256:e6360aeea578ac1d4721a0dea8088db13bbc5fd587743632e69eee33d8fda507

Observation 9f8f8f6d-c0f6-41ad-a266-038ebbde8907 · inbound

Bernini: Latent Semantic Planning for Video Diffusion cites this paper.

Bernini: Latent Semantic Planning for Video Diffusion Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model

Reference 49

Resolution
verified exact
local_arxiv, observed 2026-05-22T06:41:10.495860Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-22T06:39:47.124605Z digest=sha256:b993491637a9575ff1afa277204131a64bd01a990c936a31f32e95d818960507

Observation 52033316-44b7-4c17-9bb6-4103ef7e26a0 · inbound

Future Forcing: Future-aware Training-free KV Cache Policy for Autoregressive Video Generation cites this paper.

Future Forcing: Future-aware Training-free KV Cache Policy for Autoregressive Video Generation Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model

Reference 26

Resolution
verified exact
local_arxiv, observed 2026-06-29T08:33:15.136397Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-29T08:30:00.438334Z digest=sha256:f2e633bffcb842c10559781ccb8c135843cd0f1c27332dbcdf909592319d6cac

Observation 2003d70b-23fc-4d66-9163-b9d369efd0a3 · inbound

Veda: Scalable Video Diffusion via Distilled Sparse Attention cites this paper.

Veda: Scalable Video Diffusion via Distilled Sparse Attention Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-06-29T08:03:14.707775Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-29T07:54:13.390911Z digest=sha256:14a74fb9f23470fcd806448d4355abae369823cae42c8af0eb9396ad2b865339

Observation bfa28da5-434f-435f-acdd-876ba781297a · inbound

Diffusing in the Right Space: A Systematic Study of Latent Diffusability cites this paper.

Diffusing in the Right Space: A Systematic Study of Latent Diffusability Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model

Reference 10

Resolution
metadata mismatch
local_arxiv, observed 2026-07-02T02:36:27.627576Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-28T10:44:24.318786Z digest=sha256:df1421acabd1fd9ee69adab4b635631b44384c5385e1f8ef7e13833abc990503

Observation 5b87e6b5-7b92-4ffb-b7bc-59df73892939 · inbound

ARGUS: Stacked Multi-View Identity Mosaic Injection for Subject-Preserving Video Generation cites this paper.

ARGUS: Stacked Multi-View Identity Mosaic Injection for Subject-Preserving Video Generation Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model

Reference 46

Resolution
verified exact
local_arxiv, observed 2026-07-03T08:17:45.965420Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-27T10:46:56.871174Z digest=sha256:e0ce1676f5718db587c3ee2c1cd1ca6b7867b15f621caee6aff3ad78210f3f51

Observation be1cef6d-5305-4b0b-8ec0-2c394241452d · inbound

PhysRAG: Enhancing Physics-Awareness in Video Generation via Retrieval-Augmented Generation cites this paper.

PhysRAG: Enhancing Physics-Awareness in Video Generation via Retrieval-Augmented Generation Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model

Reference 46

Resolution
verified exact
local_arxiv, observed 2026-07-04T13:19:51.099723Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-26T05:16:53.011837Z digest=sha256:e74c5861187c174758d3f0bb7a35e83d3a65f2f4c2624059bc6d37cb2d30c912

Observation 49a440ee-1c15-421f-aa6d-ff5c44ac9f06 · inbound

Bridging Video Understanding and Generation in a Unified Framework cites this paper.

Bridging Video Understanding and Generation in a Unified Framework Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model

Reference 38

Resolution
verified exact
local_arxiv, observed 2026-07-01T10:05:40.323766Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-07-01T05:57:54.653504Z digest=sha256:7c74edf5360b14d9ffa35a2f46e62d7bc6e3eb73e5571e8cee4c63b8cf5ecde0

Observation 48e00456-ad5a-47f8-9b9c-604d58d01147 · inbound

Ink3D: Sculpting 3D Assets with Extremely Complex Textures via Video Generative Models cites this paper.

Ink3D: Sculpting 3D Assets with Extremely Complex Textures via Video Generative Models Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model

Reference 37

Resolution
metadata mismatch
local_arxiv, observed 2026-07-02T13:16:58.114308Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-07-02T13:16:16.676244Z digest=sha256:4dcd69673b2af6f6f5f96332198942e09b92f212a6ac70a552896e48f82d4424

Observation 87c8e050-6a90-45cb-9540-c7003dc306b4 · inbound

Arachne: Orchestrating Cascades for Efficient Text-to-Video Model Training cites this paper.

Arachne: Orchestrating Cascades for Efficient Text-to-Video Model Training Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model

Reference 44

Resolution
verified exact
local_arxiv, observed 2026-07-03T06:37:42.051191Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-07-03T06:27:48.935774Z digest=sha256:287eb521a15c16ae0804cf6ebff3347d079813bfcc44ca0604c81bd1b400e337

Observation a949b09d-e383-4613-8c74-1d71f7095a61 · inbound

AlayaWorld: Long-Horizon and Playable Video World Generation cites this paper.

AlayaWorld: Long-Horizon and Playable Video World Generation Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model

Reference 43

Resolution
verified exact
local_arxiv, observed 2026-07-08T10:54:49.302808Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-07-08T10:53:28.412288Z digest=sha256:068656e1b7f6bfdc8bb45573bf0295eaaff57cd0d84d7df959e6b13bbab5913e

Observation 3b0ab594-a380-4434-8d80-9a12d2a319ca · inbound

Kaleido: Algorithm-Hardware Co-Design for Video Diffusion Transformers by Exploiting Latent Space Correlations cites this paper.

Kaleido: Algorithm-Hardware Co-Design for Video Diffusion Transformers by Exploiting Latent Space Correlations Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-02T03:51:49.407043Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T03:51:49.407043Z digest=sha256:3064e49b37e2450f018b2ce8e264d41578900494af36a0692f32886aaa0bd02e

Observation 4199c132-968a-48cf-9678-9358b7234195 · inbound

DiTango: Cost-Effective Parallel Diffusion Generation with Selective Attention State Reuse cites this paper.

DiTango: Cost-Effective Parallel Diffusion Generation with Selective Attention State Reuse Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-01T22:45:27.121517Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T22:45:27.121517Z digest=sha256:15439b8c07589273980a220d1c6803ac7da564d3f4c86f3cfed004a04bf1826b

Observation 377abb5b-dc37-4e3c-acf3-546a995b061c · inbound

Multi-Dimensional Quality Assessment for AI-Generated Human-Centric Videos: Dataset and Model cites this paper.

Multi-Dimensional Quality Assessment for AI-Generated Human-Centric Videos: Dataset and Model Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-01T20:07:44.492273Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T20:07:44.492273Z digest=sha256:b72699650822ec07f0fe24d5443eeb5c28daaaeada60bf52343799f339715409

Observation da3bfc9d-02b5-4098-9c30-bb6e38287aa2 · inbound

HarmoHOI: Harmonizing Appearance and 3D Motion for Multi-view Hand-Object Interaction Synthesis cites this paper.

HarmoHOI: Harmonizing Appearance and 3D Motion for Multi-view Hand-Object Interaction Synthesis Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model

Reference 124

Resolution
unresolved
no resolver link, observed 2026-08-01T19:06:40.851302Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T19:06:40.851302Z digest=sha256:98905e9e87b8e5fe3220bae509c5ad48bf2749256d91db59e759210b22487267

Observation 3f4ba15e-323c-462c-bc24-32f845b45eeb · inbound

ABot-World-0: Infinite Interactive World Rollout on a Single Desktop GPU cites this paper.

ABot-World-0: Infinite Interactive World Rollout on a Single Desktop GPU Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-01T13:16:13.928886Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T13:16:13.928886Z digest=sha256:a282021df52cedcf6c020e024457dc34e019a3d1ca0e6cf0d1587c768302bca9

Observation fb4d8cc3-cee3-48df-a927-b53def6cf275 · inbound

VideoCoCo: Code-as-CoT for Physically-Consistent Video Generation via an Agentic Dual-Engine System cites this paper.

VideoCoCo: Code-as-CoT for Physically-Consistent Video Generation via an Agentic Dual-Engine System Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-01T08:28:03.582314Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T08:28:03.582314Z digest=sha256:625e696dd92528214a201fc0db1c1a78980e980775f7a231c58e9ef81c60c009

Observation d08a136b-85b4-4c8a-a2c8-25ec90a6b775 · inbound

Retrieval-Driven Training-Free AI-Generated Video Attribution cites this paper.

Retrieval-Driven Training-Free AI-Generated Video Attribution Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-03T16:36:47.838947Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:36:47.838947Z digest=sha256:13a426bf24f750d8959e12a0b5628efd5cf8d41debed62c0d851e9e52d74ba8e