Pith. sign in

Paper Citation Record · LEDGER

VAST 1.0: A Unified Framework for Controllable and Consistent Video Generation

As of 14 August 2026, this Paper Citation Record lists 18 of 18 outbound references and 2 inbound Pith citation observations for arXiv:2412.16677.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.16677 v1

Coverage vector

measured 18 of 18 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T10:24:49.068878Z

measured 20 of 20 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-11T23:53:47.431172Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-21T18:44:18.942934Z

Reference resolution

18 of 18 outbound references displayed

  • verified exact0
  • verified fuzzy5
  • unresolved13
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 09382561-80d6-463c-a2df-045f3f2d2440 · outbound

This paper cites Video generation models as world simulators.

VAST 1.0: A Unified Framework for Controllable and Consistent Video Generation Video generation models as world simulators

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-11T10:24:49.001121Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:24:49.001121Z digest=sha256:45e581a38a9bdaf372c728c6e2452fee8d7f422fc4ab26103309dc8de9c539c2

Observation 7df2682a-f629-4616-b242-57dfc7ddb588 · outbound

This paper cites Data-Juicer Sandbox: A Feedback-Driven Suite for Multimodal Data-Model Co-development.

VAST 1.0: A Unified Framework for Controllable and Consistent Video Generation Data-Juicer Sandbox: A Feedback-Driven Suite for Multimodal Data-Model Co-development

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-11T10:24:49.005408Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:24:49.005408Z digest=sha256:430fecca49de8ff5bfd380e73ed45ca0ac1639ce5c251225423795cd45e6eae3

Observation 8c27129b-7164-433b-9ee9-e1bca00a984a · outbound

This paper cites DreamLLM: Synergistic Multimodal Comprehension and Creation.

VAST 1.0: A Unified Framework for Controllable and Consistent Video Generation DreamLLM: Synergistic Multimodal Comprehension and Creation

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-11T10:24:49.009445Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:24:49.009445Z digest=sha256:2e73a7c0e4074fe55aa622316e1c7c7c61d008d8f65026c68ec0b9649dbdfade

Observation ba25012f-3605-4a7a-8e87-0c4f9a4b25e8 · outbound

This paper cites Scaling recti- fied flow transformers for high-resolution image synthesis.

VAST 1.0: A Unified Framework for Controllable and Consistent Video Generation Scaling recti- fied flow transformers for high-resolution image synthesis

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T10:24:49.470751Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T10:24:49.013987Z digest=sha256:45ef86d235f2c9c02305a03aab00b9b8c899b4e8cceb82ef9af643cd3fbdea09

Observation d9d54ad9-cd32-48cd-afd1-0fa606eb6d86 · outbound

This paper cites Diffit: Diffusion vision transformers for im- age generation.

VAST 1.0: A Unified Framework for Controllable and Consistent Video Generation Diffit: Diffusion vision transformers for im- age generation

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T10:24:49.458244Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T10:24:49.018340Z digest=sha256:2c84cd905c240def7de993d917311555af5d4bce7913303bdc7deddcb27ee6f5

Observation d79a8987-6c7f-489e-864a-ad9668a63011 · outbound

This paper cites SlowFast-VGen: Slow-Fast Learning for Action-Driven Long Video Generation.

VAST 1.0: A Unified Framework for Controllable and Consistent Video Generation SlowFast-VGen: Slow-Fast Learning for Action-Driven Long Video Generation

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-11T10:24:49.022079Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:24:49.022079Z digest=sha256:6eebd0b3315199046c36d806c8f61701f68d25bd3b3a32403b5c85cf95a3155b

Observation 05f61d18-0d99-49db-8404-6ef8265a1799 · outbound

This paper cites Vbench: Comprehensive bench- mark suite for video generative models.

VAST 1.0: A Unified Framework for Controllable and Consistent Video Generation Vbench: Comprehensive bench- mark suite for video generative models

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T10:24:49.445511Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T10:24:49.026013Z digest=sha256:c896b124de5fb21844da4fdcc2027096d6736b9c06e631841908129761ed7edb

Observation 5bcfc8a2-6d09-4320-acf9-f388c1ca58d4 · outbound

This paper cites HunyuanVideo: A Systematic Framework For Large Video Generative Models.

VAST 1.0: A Unified Framework for Controllable and Consistent Video Generation HunyuanVideo: A Systematic Framework For Large Video Generative Models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-11T10:24:49.029739Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:24:49.029739Z digest=sha256:b85b5d8fc75a8d8c502a222c74cf6d60915c349debbd4589484825cb89953d80

Observation 4fac470a-31af-4315-91e0-70463c397db1 · outbound

This paper cites T2v- turbo-v2: Enhancing video generation model post-training through data, reward, and conditional guidance design.arXiv preprint arXiv:2410.05677, 2024.

VAST 1.0: A Unified Framework for Controllable and Consistent Video Generation T2v- turbo-v2: Enhancing video generation model post-training through data, reward, and conditional guidance design.arXiv preprint arXiv:2410.05677, 2024

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-11T10:24:49.033826Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:24:49.033826Z digest=sha256:30587c75c591bc95ef6bc18d50cd397b48b586f346a89aab64c1795f75430703

Observation c118094e-9a73-4b41-b5ca-0254a92d7744 · outbound

This paper cites Dit-3d: Exploring plain diffusion transformers for 3d shape generation.

VAST 1.0: A Unified Framework for Controllable and Consistent Video Generation Dit-3d: Exploring plain diffusion transformers for 3d shape generation

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-11T10:24:49.037648Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:24:49.037648Z digest=sha256:d79d425b2784333c8ddab9b7b0ae4ca3175f4278e69b51b07a52ec673ff8ef58

Observation def6c258-3859-42a0-a486-30ffb2e67923 · outbound

This paper cites Scalable diffusion models with transformers.

VAST 1.0: A Unified Framework for Controllable and Consistent Video Generation Scalable diffusion models with transformers

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-11T10:24:49.043061Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:24:49.043061Z digest=sha256:31327fd8e37475595df6fdcd55e8be73dd1317b54d8deece75b05080a8bcf598

Observation c1f75ec1-89c1-45bc-bac4-a982c6823d39 · outbound

This paper cites Movie Gen: A Cast of Media Foundation Models.

VAST 1.0: A Unified Framework for Controllable and Consistent Video Generation Movie Gen: A Cast of Media Foundation Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-11T10:24:49.047094Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:24:49.047094Z digest=sha256:f56acdac6724f02ca1f7bc647e8b7ef8c85a197faa78dd62fd24f290d176ab28

Observation 2f480d80-3b41-490b-87e3-8341f7d81d84 · outbound

This paper cites DreamFusion: Text-to-3D using 2D Diffusion.

VAST 1.0: A Unified Framework for Controllable and Consistent Video Generation DreamFusion: Text-to-3D using 2D Diffusion

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-11T10:24:49.051028Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:24:49.051028Z digest=sha256:444633a05f44a9357521e94a2d9642a6d2232fffbd134183ac5ba6bf66f3f69e

Observation a65f3370-e173-4533-8e53-c5d815200eaa · outbound

This paper cites Motion-i2v: Consistent and controllable image-to-video generation with explicit motion 7 Prompt: Two Monkey Kings are locked in combat before the Heavenly Palace.

VAST 1.0: A Unified Framework for Controllable and Consistent Video Generation Motion-i2v: Consistent and controllable image-to-video generation with explicit motion 7 Prompt: Two Monkey Kings are locked in combat before the Heavenly Palace

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T10:24:49.417962Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T10:24:49.055098Z digest=sha256:00cc5971b4777ba8694529dde0bb7b7f0894782107dbad72aecdd6ce2afd6744

Observation 94d78ad9-1d78-4526-8916-c046804f7966 · outbound

This paper cites Mo- tionbooth: Motion-aware customized text-to-video genera- tion.

VAST 1.0: A Unified Framework for Controllable and Consistent Video Generation Mo- tionbooth: Motion-aware customized text-to-video genera- tion

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T10:24:49.404937Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T10:24:49.058271Z digest=sha256:179c8ea9e9fc7c6adee1cd442528312282b28056c83dfb9894242f169e81041c

Observation 7bf19f34-23ac-4f69-bd0c-09317a677ee3 · outbound

This paper cites CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer.

VAST 1.0: A Unified Framework for Controllable and Consistent Video Generation CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-11T10:24:49.061795Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:24:49.061795Z digest=sha256:a1395d8a5e4ef6560bdb40769570c307b9574cc2201e385b86350fcf1ae20262

Observation af8a5d82-04c7-4a79-9807-529b60e2f226 · outbound

This paper cites From slow bidirectional to fast causal video generators.

VAST 1.0: A Unified Framework for Controllable and Consistent Video Generation From slow bidirectional to fast causal video generators

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-11T10:24:49.065419Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:24:49.065419Z digest=sha256:9ab82e4ac3bfc212d51f6cb3b500285908331e24d35e5b6e0985305e95a081a3

Observation e085841b-7cf6-43c8-8401-05ff5f122ca6 · outbound

This paper cites Is sora a world simulator? a comprehensive survey on general world models and beyond.

VAST 1.0: A Unified Framework for Controllable and Consistent Video Generation Is sora a world simulator? a comprehensive survey on general world models and beyond

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-11T10:24:49.068878Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:24:49.068878Z digest=sha256:aacb0056529e4212e2a60f02011c27be2e6e98575e0400e80bfba163d7e2bf44

Pith citing papers

Observation b1aa08bc-0f4d-4052-994c-abd93f8ad72c · inbound

Beyond Generation: Unlocking Universal Editing via Self-Supervised Fine-Tuning cites this paper.

Beyond Generation: Unlocking Universal Editing via Self-Supervised Fine-Tuning VAST 1.0: A Unified Framework for Controllable and Consistent Video Generation

Reference 115

Resolution
unresolved
no resolver link, observed 2026-08-11T23:53:47.431172Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:53:47.431172Z digest=sha256:8fef5ba05b08322203f8cfd8a6eca0361f7b7ea3a86fec4b3fb1a3cfc04ca9f0

Observation 8b3a5cc3-1b17-4459-93f8-fa63ce348853 · inbound

Seeing What Matters: Visual Preference Policy Optimization for Visual Generation cites this paper.

Seeing What Matters: Visual Preference Policy Optimization for Visual Generation VAST 1.0: A Unified Framework for Controllable and Consistent Video Generation

Reference 38

Resolution
verified exact
arxiv_id, observed 2026-05-21T18:44:18.944678Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-21T18:42:29.327151Z digest=sha256:293c4bd5d30af51d9d47063af41d33c18ee50ad1902f964360009bac98cdcff2