Pith. sign in

Paper Citation Record · LEDGER

VAST 1.0: A Unified Framework for Controllable and Consistent Video Generation

As of 15 August 2026, this Paper Citation Record lists 18 of 18 outbound references and 2 inbound Pith citation observations for arXiv:2412.16677.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.16677 v1

Coverage vector

measured 18 of 18 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T10:24:49.068878Z

measured 20 of 20 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-11T23:53:47.431172Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-21T18:44:18.942934Z

Reference resolution

18 of 18 outbound references displayed

  • verified exact0
  • verified fuzzy5
  • unresolved13
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 09382561-80d6-463c-a2df-045f3f2d2440 · outbound

This paper cites Video generation models as world simulators.

VAST 1.0: A Unified Framework for Controllable and Consistent Video Generation Video generation models as world simulators

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-11T10:24:49.001121Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:24:49.001121Z digest=sha256:37c7cd1f48ebaa818cc82a4ddb1188f9598e5ccc2aa5340d24b51fe0cfcf15bc

Observation 7df2682a-f629-4616-b242-57dfc7ddb588 · outbound

This paper cites Data-Juicer Sandbox: A Feedback-Driven Suite for Multimodal Data-Model Co-development.

VAST 1.0: A Unified Framework for Controllable and Consistent Video Generation Data-Juicer Sandbox: A Feedback-Driven Suite for Multimodal Data-Model Co-development

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-11T10:24:49.005408Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:24:49.005408Z digest=sha256:47f1fcca7e4f39c55775c0de18101a73419c5dbd51ec35b32b265af3c1e8aff8

Observation 8c27129b-7164-433b-9ee9-e1bca00a984a · outbound

This paper cites DreamLLM: Synergistic Multimodal Comprehension and Creation.

VAST 1.0: A Unified Framework for Controllable and Consistent Video Generation DreamLLM: Synergistic Multimodal Comprehension and Creation

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-11T10:24:49.009445Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:24:49.009445Z digest=sha256:22de77e727084bc60f478ba21a5da133f027c3d4235e2649eec2af41556875b8

Observation ba25012f-3605-4a7a-8e87-0c4f9a4b25e8 · outbound

This paper cites Scaling recti- fied flow transformers for high-resolution image synthesis.

VAST 1.0: A Unified Framework for Controllable and Consistent Video Generation Scaling recti- fied flow transformers for high-resolution image synthesis

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T10:24:49.470751Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T10:24:49.013987Z digest=sha256:9db94a470f00a098211ce3f330a14e3ac2a71b89a5683468a0ad975d01aa82cf

Observation d9d54ad9-cd32-48cd-afd1-0fa606eb6d86 · outbound

This paper cites Diffit: Diffusion vision transformers for im- age generation.

VAST 1.0: A Unified Framework for Controllable and Consistent Video Generation Diffit: Diffusion vision transformers for im- age generation

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T10:24:49.458244Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T10:24:49.018340Z digest=sha256:f1c4ace26b83b60c550b2e37141fd9d3a1cf062d695a871f3b0d5643b0d6dc49

Observation d79a8987-6c7f-489e-864a-ad9668a63011 · outbound

This paper cites SlowFast-VGen: Slow-Fast Learning for Action-Driven Long Video Generation.

VAST 1.0: A Unified Framework for Controllable and Consistent Video Generation SlowFast-VGen: Slow-Fast Learning for Action-Driven Long Video Generation

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-11T10:24:49.022079Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:24:49.022079Z digest=sha256:e851d1e30a5944081df7382e0d0fd9e91fdfa37046db3cb4566cfadb6ada76ba

Observation 05f61d18-0d99-49db-8404-6ef8265a1799 · outbound

This paper cites Vbench: Comprehensive bench- mark suite for video generative models.

VAST 1.0: A Unified Framework for Controllable and Consistent Video Generation Vbench: Comprehensive bench- mark suite for video generative models

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T10:24:49.445511Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T10:24:49.026013Z digest=sha256:e09efb6b8f994a25a4de6acee12f14a3f6d090d62e7d9dba7af081f52738e1d0

Observation 5bcfc8a2-6d09-4320-acf9-f388c1ca58d4 · outbound

This paper cites HunyuanVideo: A Systematic Framework For Large Video Generative Models.

VAST 1.0: A Unified Framework for Controllable and Consistent Video Generation HunyuanVideo: A Systematic Framework For Large Video Generative Models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-11T10:24:49.029739Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:24:49.029739Z digest=sha256:4365b0206dc9a25296be3a07032bd42c030cad8a5ae86b93680610f567b3e369

Observation 4fac470a-31af-4315-91e0-70463c397db1 · outbound

This paper cites T2v- turbo-v2: Enhancing video generation model post-training through data, reward, and conditional guidance design.arXiv preprint arXiv:2410.05677, 2024.

VAST 1.0: A Unified Framework for Controllable and Consistent Video Generation T2v- turbo-v2: Enhancing video generation model post-training through data, reward, and conditional guidance design.arXiv preprint arXiv:2410.05677, 2024

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-11T10:24:49.033826Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:24:49.033826Z digest=sha256:29791af985b474963931086e02039ee1bad222bafb9988d5f355e1cbf3bf9679

Observation c118094e-9a73-4b41-b5ca-0254a92d7744 · outbound

This paper cites Dit-3d: Exploring plain diffusion transformers for 3d shape generation.

VAST 1.0: A Unified Framework for Controllable and Consistent Video Generation Dit-3d: Exploring plain diffusion transformers for 3d shape generation

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-11T10:24:49.037648Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:24:49.037648Z digest=sha256:5f05696160d61b9c55d7da186301c26535465aa101f1b8b651778431f629d44e

Observation def6c258-3859-42a0-a486-30ffb2e67923 · outbound

This paper cites Scalable diffusion models with transformers.

VAST 1.0: A Unified Framework for Controllable and Consistent Video Generation Scalable diffusion models with transformers

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-11T10:24:49.043061Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:24:49.043061Z digest=sha256:9a812d8d5a5185b8979e6eb45e5af11fe02d169a8256821f59b083cf74932d83

Observation c1f75ec1-89c1-45bc-bac4-a982c6823d39 · outbound

This paper cites Movie Gen: A Cast of Media Foundation Models.

VAST 1.0: A Unified Framework for Controllable and Consistent Video Generation Movie Gen: A Cast of Media Foundation Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-11T10:24:49.047094Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:24:49.047094Z digest=sha256:3c45844b0bb838262c983b63f6f55e57e0ef4c632ae91820482a8b6f43bc72a6

Observation 2f480d80-3b41-490b-87e3-8341f7d81d84 · outbound

This paper cites DreamFusion: Text-to-3D using 2D Diffusion.

VAST 1.0: A Unified Framework for Controllable and Consistent Video Generation DreamFusion: Text-to-3D using 2D Diffusion

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-11T10:24:49.051028Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:24:49.051028Z digest=sha256:44e985b8bf3bf88e4c49afddda5f15a9cf2eafbe0540e15a24686b7eb3a50b2e

Observation a65f3370-e173-4533-8e53-c5d815200eaa · outbound

This paper cites Motion-i2v: Consistent and controllable image-to-video generation with explicit motion 7 Prompt: Two Monkey Kings are locked in combat before the Heavenly Palace.

VAST 1.0: A Unified Framework for Controllable and Consistent Video Generation Motion-i2v: Consistent and controllable image-to-video generation with explicit motion 7 Prompt: Two Monkey Kings are locked in combat before the Heavenly Palace

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T10:24:49.417962Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T10:24:49.055098Z digest=sha256:c194ef56c21c0c9440fdf9dc86a141bc86ecab845d94b0f9840bac931ceebcc3

Observation 94d78ad9-1d78-4526-8916-c046804f7966 · outbound

This paper cites Mo- tionbooth: Motion-aware customized text-to-video genera- tion.

VAST 1.0: A Unified Framework for Controllable and Consistent Video Generation Mo- tionbooth: Motion-aware customized text-to-video genera- tion

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T10:24:49.404937Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T10:24:49.058271Z digest=sha256:074f238dee2eaa4fee4bc72f94bf5f68576df560da1a28fe6a757e92b0168b6d

Observation 7bf19f34-23ac-4f69-bd0c-09317a677ee3 · outbound

This paper cites CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer.

VAST 1.0: A Unified Framework for Controllable and Consistent Video Generation CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-11T10:24:49.061795Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:24:49.061795Z digest=sha256:9c009a9754c2fa722cbf0b007dc179d3cb3b93f09d3a07b412fa9b9e387dc0a5

Observation af8a5d82-04c7-4a79-9807-529b60e2f226 · outbound

This paper cites From slow bidirectional to fast causal video generators.

VAST 1.0: A Unified Framework for Controllable and Consistent Video Generation From slow bidirectional to fast causal video generators

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-11T10:24:49.065419Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:24:49.065419Z digest=sha256:2387fe37079e04b3739fc5306d3c5172adcf516a2e1887407133d2e3eef2039f

Observation e085841b-7cf6-43c8-8401-05ff5f122ca6 · outbound

This paper cites Is sora a world simulator? a comprehensive survey on general world models and beyond.

VAST 1.0: A Unified Framework for Controllable and Consistent Video Generation Is sora a world simulator? a comprehensive survey on general world models and beyond

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-11T10:24:49.068878Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:24:49.068878Z digest=sha256:f0e3f949d9702151037f451e8caaa8e55d34210a1110ddf90575771621f0bf84

Pith citing papers

Observation b1aa08bc-0f4d-4052-994c-abd93f8ad72c · inbound

Beyond Generation: Unlocking Universal Editing via Self-Supervised Fine-Tuning cites this paper.

Beyond Generation: Unlocking Universal Editing via Self-Supervised Fine-Tuning VAST 1.0: A Unified Framework for Controllable and Consistent Video Generation

Reference 115

Resolution
unresolved
no resolver link, observed 2026-08-11T23:53:47.431172Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:53:47.431172Z digest=sha256:a05236ef0f1074d923851b3ecc4bbb39a108cb0dec1655ecdb28fde026652ef1

Observation 8b3a5cc3-1b17-4459-93f8-fa63ce348853 · inbound

Seeing What Matters: Visual Preference Policy Optimization for Visual Generation cites this paper.

Seeing What Matters: Visual Preference Policy Optimization for Visual Generation VAST 1.0: A Unified Framework for Controllable and Consistent Video Generation

Reference 38

Resolution
verified exact
arxiv_id, observed 2026-05-21T18:44:18.944678Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-21T18:42:29.327151Z digest=sha256:992c0e133ab0cfe58251e5e4a014bf2e93bc19762b3511c68a2004d37068192d