Pith. sign in

Paper Citation Record · LEDGER

VidGen-1M: A Large-Scale Dataset for Text-to-video Generation

As of 22 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 27 inbound Pith citation observations for arXiv:2408.02629.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2408.02629 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 27 of 27 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00

measured 27 of 27 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-14T04:21:49.967350Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-09T07:56:04.692157Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation d3ea8e14-b555-4cb7-8a88-adabe49cab4e · inbound

Towards World Simulator: Crafting Physical Commonsense-Based Benchmark for Video Generation cites this paper.

Towards World Simulator: Crafting Physical Commonsense-Based Benchmark for Video Generation VidGen-1M: A Large-Scale Dataset for Text-to-video Generation

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-05-18T14:40:00.032419Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-05-18T14:39:59.870039Z digest=sha256:d3ae32bdad9df81ab7dc9feb52960df922e1038d0e0745c2e0acf33f35539923

Observation d3c18b2b-2818-4fc1-9ad7-c5af5c13b606 · inbound

EgoVid-5M: A Large-Scale Video-Action Dataset for Egocentric Video Generation cites this paper.

EgoVid-5M: A Large-Scale Video-Action Dataset for Egocentric Video Generation VidGen-1M: A Large-Scale Dataset for Text-to-video Generation

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-12T21:41:59.627308Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:41:59.627308Z digest=sha256:2846567c3afd7ff0b2126c2c706f638677cf75071692f6a6c8c318d967ee459b

Observation 622021ba-948d-41b7-b30c-868fd9f0e87b · inbound

Individual Content and Motion Dynamics Preserved Pruning for Video Diffusion Models cites this paper.

Individual Content and Motion Dynamics Preserved Pruning for Video Diffusion Models VidGen-1M: A Large-Scale Dataset for Text-to-video Generation

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-12T11:20:37.961513Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:20:37.961513Z digest=sha256:bc6e101c0414900d5d1a2d3dc10484a0b434f0e830c3a03d20c7d3a03a37880b

Observation 5aa572ee-3d91-417c-9f34-4d0026793aa7 · inbound

BlobGEN-Vid: Compositional Text-to-Video Generation with Blob Video Representations cites this paper.

BlobGEN-Vid: Compositional Text-to-Video Generation with Blob Video Representations VidGen-1M: A Large-Scale Dataset for Text-to-video Generation

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-10T20:41:46.844463Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:41:46.844463Z digest=sha256:ab7a8ee321707c6c75e1c8a7d4d50f3b0995fce43ca7ec5b3bcc98fb2bfcbe64

Observation 67f6d166-4f8b-4b80-aed4-917be129f93e · inbound

On-device Sora: Enabling Training-Free Diffusion-based Text-to-Video Generation for Mobile Devices cites this paper.

On-device Sora: Enabling Training-Free Diffusion-based Text-to-Video Generation for Mobile Devices VidGen-1M: A Large-Scale Dataset for Text-to-video Generation

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-09T10:46:29.559618Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T10:46:29.559618Z digest=sha256:3ba5174e0300c58c79b8e1403b633a92ed620fc5838e1ea4cb12adb8eb392d75

Observation e562b6aa-ddbf-4d8b-8895-fcd721069b93 · inbound

TUNA: Comprehensive Fine-grained Temporal Understanding Evaluation on Dense Dynamic Videos cites this paper.

TUNA: Comprehensive Fine-grained Temporal Understanding Evaluation on Dense Dynamic Videos VidGen-1M: A Large-Scale Dataset for Text-to-video Generation

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T14:03:03.857463Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:03:03.857463Z digest=sha256:675ff1c3eeacf0abac31479b0e2dee21888215ed8977bfdd77d3499ca0c9c84a

Observation 01bdbe61-9ce9-47c0-a549-2bbcf545fcdb · inbound

VCapsBench: A Large-scale Fine-grained Benchmark for Video Caption Quality Evaluation cites this paper.

VCapsBench: A Large-scale Fine-grained Benchmark for Video Caption Quality Evaluation VidGen-1M: A Large-Scale Dataset for Text-to-video Generation

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T12:48:06.335779Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:48:06.335779Z digest=sha256:2b21988310337bd51de1e7f0686f8ff4371b71f6a51183702e6096456a9ed5fc

Observation 021bbf87-c886-4d99-b10e-ebaea25b0b1e · inbound

UltraVideo: High-Quality UHD Video Dataset with Comprehensive Captions cites this paper.

UltraVideo: High-Quality UHD Video Dataset with Comprehensive Captions VidGen-1M: A Large-Scale Dataset for Text-to-video Generation

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T00:30:58.490871Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:30:58.490871Z digest=sha256:e0f90843032db6a56e9a3b19f4293d6f6d483f405fad97c98e2f7fbbf20e92e8

Observation f858776c-a47b-478c-b409-976b86053609 · inbound

MiCo: Multi-image Contrast for Reinforcement Visual Reasoning cites this paper.

MiCo: Multi-image Contrast for Reinforcement Visual Reasoning VidGen-1M: A Large-Scale Dataset for Text-to-video Generation

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-06T22:10:21.932155Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:10:21.932155Z digest=sha256:164645f33acbd647c8295d0596bfff6e9be25508816d061dd0d0a5aeec0ba15c

Observation ccb936a4-6747-4788-a868-d7a8dd1ba52a · inbound

FreeLong++: Training-Free Long Video Generation via Multi-band SpectralFusion cites this paper.

FreeLong++: Training-Free Long Video Generation via Multi-band SpectralFusion VidGen-1M: A Large-Scale Dataset for Text-to-video Generation

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T21:28:17.589175Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:28:17.589175Z digest=sha256:3bfe22af1aa91b5d87b40dfa383dbda421d5310c2e0a91c7a4a5b031287c3805

Observation 390de76e-87ca-478f-acca-2019bac7a998 · inbound

PhyDetEx: Detecting and Explaining the Physical Plausibility of T2V Models cites this paper.

PhyDetEx: Detecting and Explaining the Physical Plausibility of T2V Models VidGen-1M: A Large-Scale Dataset for Text-to-video Generation

Reference 40

Resolution
verified exact
arxiv_id, observed 2026-05-21T18:00:27.203120Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-21T17:57:57.263574Z digest=sha256:9dde0af20befb51dcad4b9e1d740d31e1d6d63773124a20dad3c9fab296cbff6

Observation e9a2f9df-bd54-4244-9eed-4e6bf4ae1808 · inbound

Mind the Generative Details: Direct Localized Detail Preference Optimization for Video Diffusion Models cites this paper.

Mind the Generative Details: Direct Localized Detail Preference Optimization for Video Diffusion Models VidGen-1M: A Large-Scale Dataset for Text-to-video Generation

Reference 56

Resolution
verified exact
arxiv_id, observed 2026-05-16T16:28:05.580957Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-16T16:25:03.743594Z digest=sha256:76c9a74f48cf396dd374ab4a79591b669d7202530478d67a8b00928a01d768b6

Observation c9616a39-1e50-41c6-94d1-c11ac9fe2134 · inbound

Mind the Generative Details: Direct Localized Detail Preference Optimization for Video Diffusion Models cites this paper.

Mind the Generative Details: Direct Localized Detail Preference Optimization for Video Diffusion Models VidGen-1M: A Large-Scale Dataset for Text-to-video Generation

Reference 56

Resolution
verified exact
arxiv_id, observed 2026-05-21T16:04:14.709987Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-21T16:01:52.150950Z digest=sha256:55efd935056521f27e89fe554522ee13325ce13c86c402c7b8f123101acecd88

Observation c77cf558-2b39-49cc-b1f4-fae8875e98cb · inbound

VDCook:DIY video data cook your MLLMs cites this paper.

VDCook:DIY video data cook your MLLMs VidGen-1M: A Large-Scale Dataset for Text-to-video Generation

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-15T16:56:19.221954Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-15T16:52:36.882218Z digest=sha256:788643a711a5b4dc7a99c3c5ca32cefd3459243e502a3b74f101b159c4fd93e2

Observation 08fdc825-9988-48cb-86a3-653e8224fdc7 · inbound

Can Video Diffusion Models Predict Past Frames? Bidirectional Cycle Consistency for Reversible Interpolation cites this paper.

Can Video Diffusion Models Predict Past Frames? Bidirectional Cycle Consistency for Reversible Interpolation VidGen-1M: A Large-Scale Dataset for Text-to-video Generation

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-02T16:56:33.961865Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T16:56:33.961865Z digest=sha256:6c138c3777a3260d65efc7da09bd6e8116024cb28d9749c65129c1bd33f2d03b

Observation c20ae787-89ea-405c-92c4-52f261aa6189 · inbound

PhyEdit: Towards Real-World Object Manipulation via Physically-Grounded Image Editing cites this paper.

PhyEdit: Towards Real-World Object Manipulation via Physically-Grounded Image Editing VidGen-1M: A Large-Scale Dataset for Text-to-video Generation

Reference 47

Resolution
verified exact
arxiv_id, observed 2026-05-10T23:45:50.458209Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-10T18:54:47.451103Z digest=sha256:3d2ee9600c6b82589c5e1df7d8c70f6caae541e117788f3aee14c75692e03fa6

Observation e1562e97-2246-4edc-921f-17f561d6d65e · inbound

OmniShotCut: Holistic Relational Shot Boundary Detection with Shot-Query Transformer cites this paper.

OmniShotCut: Holistic Relational Shot Boundary Detection with Shot-Query Transformer VidGen-1M: A Large-Scale Dataset for Text-to-video Generation

Reference 32

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T21:46:49.069232Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-08T04:15:20.045060Z digest=sha256:a6669f434fb0243339aecb54dbabda205622eda3b43ea6d13298d7a4434a686f

Observation cdfcdd4e-e450-4e2a-a17b-86b549ad0622 · inbound

Variance Reduction for Expectations with Diffusion Teachers cites this paper.

Variance Reduction for Expectations with Diffusion Teachers VidGen-1M: A Large-Scale Dataset for Text-to-video Generation

Reference 79

Resolution
verified exact
arxiv_id, observed 2026-05-21T04:53:57.919961Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-21T04:51:57.800006Z digest=sha256:6e35216cbb55eaea2af8a47a9bc64eca34954024632ac18995f9f480744e82a7

Observation 3e10c653-0cca-4927-a583-1d6cceba155a · inbound

Variance Reduction for Expectations with Diffusion Teachers cites this paper.

Variance Reduction for Expectations with Diffusion Teachers VidGen-1M: A Large-Scale Dataset for Text-to-video Generation

Reference 79

Resolution
verified exact
arxiv_id, observed 2026-05-25T05:45:24.685953Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-25T05:41:10.076206Z digest=sha256:78a3d12de0171ca78e1d5573f3db65f00f01b6da3958ac73c6d7421232168d11

Observation 1fa789f5-e1b1-4950-8b96-b313095709a1 · inbound

Toward Native Multimodal Modeling: A Roadmap cites this paper.

Toward Native Multimodal Modeling: A Roadmap VidGen-1M: A Large-Scale Dataset for Text-to-video Generation

Reference 115

Resolution
verified exact
arxiv_id, observed 2026-06-29T23:04:01.932824Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-29T22:58:38.610609Z digest=sha256:a083959132ee85725c456217942eedf2bd9d06ad044f5b551b0be849f4053764

Observation f9d4710d-4d57-43be-8766-893118d790e6 · inbound

JAVEDIT: Joint Audio-Visual Instruction-Guided Video Editing with Agentic Data Curation cites this paper.

JAVEDIT: Joint Audio-Visual Instruction-Guided Video Editing with Agentic Data Curation VidGen-1M: A Large-Scale Dataset for Text-to-video Generation

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-07-02T02:16:26.776236Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-28T11:04:49.366127Z digest=sha256:a8294610a2231a37a5a8516a8df535f8ef2e0eed0247028647d36915225be48f

Observation 0f8ec3eb-70b6-4884-b4e6-72129c28253c · inbound

Physics-Informed Video Generation via Mixture-of-Experts Latent Alignment cites this paper.

Physics-Informed Video Generation via Mixture-of-Experts Latent Alignment VidGen-1M: A Large-Scale Dataset for Text-to-video Generation

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-07-02T07:16:44.728428Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-28T07:02:37.291472Z digest=sha256:2e138e189a26082576517325eb28b34b8b636a45685f43e1792481648cd4ffb9

Observation a6023803-aa52-45f6-997c-ca561d221cb9 · inbound

CineDance: Towards Next-Generation Multi-Shot Long-Form Cinematic Audio-Video Generation cites this paper.

CineDance: Towards Next-Generation Multi-Shot Long-Form Cinematic Audio-Video Generation VidGen-1M: A Large-Scale Dataset for Text-to-video Generation

Reference 62

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T00:07:28.154541Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-27T17:30:25.371658Z digest=sha256:928c979ddd96902f5f9f69d2307d48f6c7cb0bc76c912add56b3ab0861fb156a

Observation dcd6ee96-b556-420b-b064-51f048eb7cce · inbound

Kairos: A Regret-Aware Native World-Action Model Stack for Physical AI cites this paper.

Kairos: A Regret-Aware Native World-Action Model Stack for Physical AI VidGen-1M: A Large-Scale Dataset for Text-to-video Generation

Reference 42

Resolution
verified exact
arxiv_id, observed 2026-07-03T17:28:44.842107Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-27T04:02:53.110012Z digest=sha256:a6041a0498071766ec77b958ea1091446a0b6c4c4d5c5285d901efd990acc219

Observation 5f133fd3-a200-47b8-b965-12e576e6ac11 · inbound

Infinite Worlds with Versatile Interactions cites this paper.

Infinite Worlds with Versatile Interactions VidGen-1M: A Large-Scale Dataset for Text-to-video Generation

Reference 42

Resolution
verified exact
local_arxiv, observed 2026-07-09T07:56:04.693407Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-07-09T07:51:36.802801Z digest=sha256:64585cd76faa3010a401ae2ff637d7e22398d6ac4d4a7980857cb451b31cbd5e

Observation 500fa9be-3d54-4470-8bef-a7b993569353 · inbound

Group-of-Latents: Perceptual Video Compression at Extreme Bitrates via Masked Latent Generative Modeling cites this paper.

Group-of-Latents: Perceptual Video Compression at Extreme Bitrates via Masked Latent Generative Modeling VidGen-1M: A Large-Scale Dataset for Text-to-video Generation

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-01T14:30:46.262409Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T14:30:46.262409Z digest=sha256:e08d860cb06e5c7075d073017f668bb8105d7eeae959b765c89d8e3bf9f10240

Observation 6bbec479-b087-4d0d-baa6-b432d9f7a415 · inbound

Sekai2: From World Exploration to Interactive World Modeling cites this paper.

Sekai2: From World Exploration to Interactive World Modeling VidGen-1M: A Large-Scale Dataset for Text-to-video Generation

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-14T04:21:49.967350Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:21:49.967350Z digest=sha256:05a5b06bc3b224d5888a187882496219a748bb81287a506e4de7a70ddb69fe36